Vector databases power AI retrieval and semantic search. How they work, when to use one and the security and cost questions IT teams should ask.
Generative AI applications increasingly rely on retrieval: finding the most relevant internal documents and passing them to a language model as context. Vector search is the technology that makes this work.
What is a vector?
An embedding model converts text, images or other content into a list of numbers — a vector — that captures its meaning. Items with similar meanings produce vectors that are close together, so a search for “employee leave policy” can find a document titled “time-off guidelines” even with no shared keywords.
What a vector database does
A vector database stores these embeddings and finds the nearest matches quickly using approximate nearest-neighbour indexes. Many also store metadata so results can be filtered — for example by department, document type or access rights.
Dedicated database or existing platform?
Purpose-built vector databases offer advanced indexing and scale. Many established databases and search engines now also support vector search, which can be simpler if your data already lives there. Choose based on scale, latency needs, operational skills and how well each option enforces access control.
Operational questions
- Security: can users only retrieve documents they are allowed to see?
- Freshness: how quickly are new and changed documents re-embedded?
- Quality: how will you evaluate whether retrieved results are relevant?
- Cost: what are the storage, indexing and embedding costs at your volume?
- Portability: if you change embedding models, how will you rebuild indexes?
5 best practices for production vector search
- Chunk content thoughtfully. Split documents into passages that preserve meaning, such as sections or paragraphs, and keep titles and headings with each chunk for context.
- Store rich metadata. Source, owner, date, classification and access groups allow filtering and help enforce permissions at query time.
- Combine semantic and keyword search. Hybrid retrieval often improves accuracy for product codes, names and acronyms that embeddings handle poorly.
- Evaluate retrieval quality. Build a test set of real questions with known good answers and measure whether the right passages are retrieved before tuning prompts.
- Automate updates. Re-embed changed documents and remove deleted ones automatically so answers do not rely on stale or withdrawn content.
Where vector search fits in an AI architecture
In a retrieval-augmented generation (RAG) design, a user question is converted to an embedding, the most relevant passages are retrieved and the language model generates an answer grounded in those passages. This reduces hallucinations and lets the model use current, internal knowledge without retraining.
Common mistakes to avoid
- Indexing everything without considering who should see it.
- Choosing a large, expensive embedding model without testing whether a smaller one performs as well.
- Ignoring index rebuild time when planning model upgrades.
- Measuring only response speed instead of answer quality.
Frequently asked questions
Do we need a separate vector database?
Not always. If your data already lives in a database or search engine with vector support, start there and move only if scale or features require it.
How much data can vector search handle?
Modern systems handle millions to billions of vectors, but memory, index type and filtering needs strongly affect cost and latency.
A 90-day action plan
Days 1 to 30: select one knowledge domain, such as HR policies or IT support articles, collect the source documents and write fifty real questions with expected answers.
Days 31 to 60: build a retrieval prototype, test several chunk sizes and embedding models against your question set, and add permission filtering based on existing access groups.
Days 61 to 90: pilot with a small group of users, collect feedback on answer quality, measure cost per query and plan automated content refresh.
Questions to ask vendors
- How do you enforce document-level permissions at query time?
- Which index types are supported, and how do they trade recall against speed?
- How are updates and deletions handled without full rebuilds?
- Is hybrid keyword and semantic search supported natively?
- Where is information stored and processed, and is it used to train models?
Key terms explained
- Embedding: a numeric representation of content that captures meaning.
- Approximate nearest neighbour: an indexing method that finds similar items quickly with small accuracy trade-offs.
- Retrieval-augmented generation: supplying retrieved passages to a language model to ground its answer.
- Chunking: splitting documents into passages before indexing.
- Recall: the share of relevant results that a search actually returns.
The bottom line
Semantic retrieval has become a core building block for AI assistants and enterprise search. The technology choice matters less than the practices around it: thoughtful chunking, rich metadata, permission filtering, hybrid search, systematic evaluation and automated updates. Start with one knowledge domain and a realistic set of test questions, measure answer quality and cost, and expand once results are reliable. Strong governance keeps AI answers accurate and secure.
Further reading on vector databases
For authoritative, vendor-neutral guidance on vector databases, see the open-source pgvector project. You can also browse our free whitepapers.

