Vector Database Sizing Calculator
Memory and storage for embeddings: HNSW, IVF or PQ, with every formula shown.
Your data
Index
Copies, growth and units
Memory and disk
| Part | Formula | Size | Where |
|---|
A year of growth
| Month | Vectors | Memory | Disk |
|---|
Cost of creating the embeddings (optional)
About the Vector Database Sizing Calculator
Plan the memory and storage of a retrieval-augmented generation (RAG) index before you buy the servers. Enter how many chunks you will embed, the embedding dimensions and precision, the index type and its settings — HNSW with M, or IVF with its lists and product-quantization code size — the metadata stored with each vector, the number of copies and how fast the data grows.
The calculator shows the raw vectors, what the index adds, the memory an in-memory index needs, the disk it takes and a year of growth, each line with its formula and your numbers. You can also add the cost of creating the embeddings from the tokens per chunk and a price you enter, and read the trade-offs of quantization and shorter embeddings.
How to use it
- Enter the number of documents and chunks per document, or the number of chunks directly.
- Enter the dimensions of your embedding model and the precision you store: float32, float16, int8 or binary.
- Choose the index: Flat for exact search on small sets, HNSW with its M, IVF-Flat with its lists or IVF-PQ with lists and a code size in bytes.
- Add the metadata bytes per vector (text, ids, tags you store with it), the copies you run and the growth per month. Pick binary (GiB) or decimal (GB) units.
- Read memory and disk per copy and in total, check the formulas, and look at the month-by-month table. Copy the summary or download the table as CSV.
- Optionally enter the tokens per chunk and your embedding model’s price per 1M tokens to see what creating the embeddings costs.
Examples
float32 · HNSW, M = 16 · 8-byte ids · 200 bytes of metadata per chunk · 1 copy
Vectors 6.14 GB · graph 128.0 MB · ids 8.00 MB · memory 6.28 GB · disk 6.48 GB
Faiss’s formula for HNSW is (d × 4 + M × 2 × 4) bytes per vector: 6,144 + 128 bytes here.
4,000 lists · 96-byte codes · full vectors kept on disk for re-ranking
Memory 130.1 MB (codes 96.0 MB, centroids 24.58 MB, codebooks 1.57 MB, ids 8.00 MB) · disk 6.47 GB
Each vector shrinks from 6,144 bytes to a 96-byte code, 64 times smaller, at some cost in recall.
1 million chunks · 1,536 dimensions · binary · HNSW, M = 16
Vectors 192.0 MB · graph 128.0 MB · ids 8.00 MB · memory 328.0 MB
Common uses
- Choosing a server or a managed plan with enough memory for a RAG index.
- Comparing HNSW with IVF-PQ, or float32 with float16, int8 and binary vectors, before you build.
- Planning how much memory and disk a year of new documents needs.
- Estimating the one-off and monthly cost of embedding your documents.
The formulas
- Vectors: chunks × dimensions × bytes per dimension (float32 4, float16 2, int8 1, binary ⅛). pgvector’s README gives the same sizes for its types: 4 × d + 8 bytes for vector, 2 × d + 8 for halfvec and d ÷ 8 + 8 for bit.
- HNSW: “(d * 4 + M * 2 * 4) bytes per vector” (Faiss guidelines): the vector plus M × 2 links of 4 bytes on the bottom layer. hnswlib puts the graph at roughly M × 8–10 bytes per element, the upper layers adding the rest.
- IVF: each vector also stores an 8-byte id (Faiss indexes: IVF-Flat 4 × d + 8, IVF-PQ code size + 8), and the index keeps one float32 centroid per list.
- PQ: with 8-bit codes the code size in bytes equals the number of sub-vectors, and each sub-vector has 256 centroids: 256 × d floats of codebooks in all.
- Memory adds up the parts an in-memory index needs; disk adds the metadata and, for IVF-PQ, the full vectors kept for re-ranking. Both are multiplied by the number of copies.
Choosing an index
Flat compares the query with every vector: exact, and fine for tens of thousands of vectors. HNSW is a graph that finds neighbours fast at the cost of M × 8 bytes per vector; M 12–48 suits most data, and ef_construction and ef_search trade build and query time for recall without changing memory (hnswlib). IVF-Flat splits the vectors into lists and searches only a few of them. Faiss’s guidelines suggest 4√N to 16√N lists below a million vectors, then 65,536 lists up to 10 million, 262,144 up to 100 million and 1,048,576 up to a billion, trained on 30 to 256 vectors per list; pgvector suggests rows ÷ 1000 up to a million rows and √rows above. IVF-PQ also compresses each vector into a short code, which saves the most memory and needs re-ranking on the full vectors for the best recall.
Smaller vectors: precision and dimensions
Lower precision shrinks memory in proportion: float16 halves it, int8 quarters it and binary codes are a thirty-second of float32. How much recall you lose depends on the model and the data, so test it on your own queries. Some embedding models are trained so that their vectors can be shortened: OpenAI’s embeddings guide reports that a text-embedding-3-large embedding shortened to 256 dimensions still beats a full 1,536-dimension text-embedding-ada-002 one on the MTEB benchmark. pgvector indexes up to 2,000 dimensions for vector, 4,000 for halfvec and 64,000 for bit columns.
Limitations
- An estimate from the published per-vector formulas: every database adds its own overhead (segment files, write-ahead logs, deleted vectors waiting for compaction, page alignment), so leave headroom.
- Building an index needs extra memory for a while (pgvector, for example, builds HNSW fastest when the graph fits in maintenance_work_mem); that peak is not included.
- HNSW is counted with its bottom layer of links (the Faiss formula); hnswlib’s M × 8–10 bytes includes the upper layers, a few per cent more.
- Recall, latency and the prices of managed vector databases are not estimated.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
How much memory do 1 million OpenAI embeddings need?
text-embedding-3-small returns 1,536 dimensions by default: 1 million float32 vectors take 6.14 GB (5.72 GiB), and an HNSW index with M = 16 adds 128 MB of links plus the ids. Stored as float16 they take half that, and as binary codes 192 MB.
Do ef_construction and ef_search change the memory?
No. They set how many candidates HNSW looks at while building and searching, so they change build time, query time and recall. Memory depends on the vectors and M.
GiB or GB?
Memory is usually sold in binary units (1 GiB = 1,073,741,824 bytes) and disks in decimal ones (1 GB = 1,000,000,000 bytes). Choose the units that match the plan or server you compare with.
What should I enter as metadata bytes?
The average size of everything stored with each vector: the chunk text if the database keeps it, document ids, titles, tags and dates. A chunk of 500 English words is roughly 3 KB of text.
Is anything uploaded?
No. The calculator runs in your browser and works offline once the page has loaded.