Skip to main content

One doc tagged with "Vision"

View all tags

Multi-Modal RAG

Retrieval over documents that mix text and images — extract the text, have a vision model describe each image as a caption, embed both into one vector store, and answer from whichever is relevant.