Exam Room · Advanced Generative AI Developer

Choosing an Embedding Dimension and Its Storage Cost

· 26 min read

Generative AI Development · part of The Exam Room

The situation

A knowledge-base team is building retrieval-augmented answering on Amazon Bedrock. They have roughly ten million passages today, growing toward maybe forty million as more document sources come online. Each passage is embedded and stored in a vector index that a retrieval step queries on every question, and the answer quality depends on the top handful of passages coming back relevant.

They started on Amazon Titan Text Embeddings v2 at its default 1024 dimensions because that was the number in the first tutorial they read. The index is already tens of gigabytes, the managed vector store bill is the fastest-growing line in the account, and query latency at the p99 is starting to be noticeable in the chat experience. Someone asked the obvious question: Titan v2 emits 512 or 256 dimensions on request, so could they halve or quarter the footprint, and how far would answer quality fall?

Re-embedding forty million passages twice by trial and error is a week nobody has. A smaller index that returns worse passages raises no error to say so, either. The decision underneath is the one every embedding project reaches: how many dimensions, given this corpus size, this quality bar, and this storage and latency budget.

What actually matters

An embedding dimension is how many numbers each vector carries, and every one of those numbers is stored and compared for every vector in the index. Dimension multiplies against the number of vectors and against the bytes each number takes. Doubling the dimension roughly doubles the raw vector bytes, the memory the index holds, and the arithmetic each similarity comparison performs. On a ten-passage toy index none of this matters. On tens of millions of passages it is the dominant term in the bill.

More dimensions can represent more nuance. A higher-dimensional space has more room to separate subtly different meanings, so on hard, semantically dense corpora a larger embedding often retrieves better. The catch is that the relationship is not linear and it is not guaranteed. Past the point where the extra dimensions stop capturing distinctions your queries depend on, storage and latency keep climbing while retrieval quality flattens out. Bigger is a hypothesis to test, not a rule to assume.

A lower dimension reduces every axis at once: less storage, a smaller in-memory index, faster distance computations, and often lower latency. Some fidelity goes with it, and how much depends on your data and your quality bar. Some corpora lose almost nothing in the drop from 1024 to 512; others fall off a cliff. The only way to know which one you have is to measure retrieval quality at each dimension against a set of real queries with known-good answers, using something like Recall (retrieval)The share of genuinely relevant passages a search actually returns – what you lose when you retrieve fewer chunks. or a downstream answer-quality score, rather than reasoning about it in the abstract.

Titan Text Embeddings v2 (amazon.titan-embed-text-v2:0) makes this trade explicit. Its dimensions request field accepts 1024, 512, or 256, and defaults to 1024. The shorter vectors are a supported output of the model rather than a truncation you perform yourself, and AWS’s own benchmarking reported about 3.24% accuracy loss at 256 dimensions against a four-fold reduction in size. Configurability is a property of the model, not of embeddings in general. Titan Embeddings G1 - Text (amazon.titan-embed-text-v1) emits a fixed 1536 floats. Cohere Embed English v3 and Embed Multilingual v3 are fixed at 1024, while Cohere Embed v4 accepts 256, 512, 1024, or 1536 and defaults to 1536.

Two constraints sit underneath all of this and break the index silently if you get them wrong. The distance metric and the model have to match what the index was built for. An index configured for cosine similarity expects vectors compared by angle; one configured for Euclidean or dot-product expects something else, and querying with the wrong metric returns plausible-looking nonsense. The same holds for the model and dimension: every vector in one index must come from the same model at the same dimension, because vectors from different models or different sizes do not live in a comparable space. Re-embedding at a new dimension means rebuilding the index, not mixing sizes in place.

Normalisation sits alongside the metric. Titan v2 takes a normalize flag that defaults to true, returning unit-length vectors, and that default is usually the one you want: with normalised vectors, cosine similarity and dot product rank results identically, and cosine is what most retrieval setups assume. Set it to false and dot-product magnitudes reflect vector length as well as direction, which changes the ranking. Normalise consistently and pair it with a matching metric, across indexing and querying alike.

What we’ll filter on

  1. Corpus size, how many vectors the index holds now and at projected growth, since dimension multiplies against every one of them.
  2. Retrieval-quality bar, the recall@k or answer-quality floor the application needs, measured on real queries rather than assumed.
  3. Storage and memory footprint, the raw vector bytes plus index overhead the budget can carry.
  4. Query latency, whether smaller vectors and a smaller index keep p99 inside the experience’s budget.
  5. Model support, whether the chosen model actually offers configurable dimensions, and which sizes.
  6. Metric and normalisation fit, whether the distance metric and normalisation match across the model, the index, and the query path.

The landscape

1024 dimensions (Titan v2 default). The most nuance the model offers and the safest starting point for quality, since nothing is compressed. It is also the heaviest on every axis: largest vectors, largest index, most arithmetic per comparison. Use it as the baseline you measure everything else against, and stay there if your data needs the fidelity and the budget allows it.

512 dimensions. Half the raw storage and roughly half the per-comparison work of 1024, with a fidelity drop that is often small on well-behaved corpora. This is frequently the sweet spot for large indexes, where the footprint saving is real and the recall drop, if any, sits inside the quality bar. Measure it against the 1024 baseline first.

256 dimensions. A quarter of the 1024 storage and the fastest to search. The fidelity drop is larger and more corpus-dependent, so it fits a corpus that is huge and cost-sensitive, with queries that are reasonably distinguishable and measurement confirming recall still clears the bar.

Other models, other ladders. The dimension list comes with the model. Titan Embeddings G1 - Text is fixed at 1536 floats, and Cohere Embed English v3 and Multilingual v3 are fixed at 1024, so on any of them the only way to change footprint is to change models and re-embed. Cohere Embed v4 has a ladder of its own, 256, 512, 1024 and 1536, and handles images alongside text. Switching models rebuilds the index and restarts the quality measurement, so it is a larger move than turning the dial on the model already in place.

Quantisation, an orthogonal lever. Separate from dimension is how many bytes each number takes. Vectors are commonly stored as 32-bit floats, 4 bytes each. Titan v2 returns binary embeddings directly through its embeddingTypes field, Cohere Embed v4 offers int8, uint8, binary and ubinary, and a Bedrock knowledge base takes an embeddingDataType of FLOAT32 or BINARY when you create it. Binary at 1024 dimensions is 128 bytes a vector against 4,096, with its own quality trade to measure. Dimension and quantisation stack, so run them as two separate experiments and attribute the quality change correctly.

Evaluation

Side by side

Option Relative storage Search speed Retrieval fidelity Tunable Best when
Titan v2, 1024 Highest (baseline) Slowest Highest ✓ Dense corpus, quality-led, budget allows it
Titan v2, 512 ~½ of 1024 Faster Usually close to 1024 ✓ Large index, footprint matters, small recall drop
Titan v2, 256 ~¼ of 1024 Fastest Lower, corpus-dependent ✓ Huge, cost-sensitive index, quality still clears bar
Titan G1, fixed 1536 Highest, fixed Slowest High ✗ Older model; no dimension lever
Cohere Embed v4, 256-1536 Scales with chosen size Scales with size Varies by size ✓ Different model; also embeds images
Quantised storage (int8 / binary) Cuts bytes-per-value Faster Trade to measure ✓ Stacks on any dimension to cut footprint further

Reading the table against the ten-to-forty-million-passage index: 1024 is the quality baseline to beat, 512 is the first serious candidate for halving the footprint, and 256 is in play if measurement says the corpus tolerates it. Quantisation is a second, independent saving to layer on once the dimension is settled. Changing models is a bigger job than either.

The solution

Start with the storage arithmetic, which shows how much the choice is worth. The raw vector footprint is vectors times dimension times bytes-per-value. At ten million passages, 1024 dimensions, and 4-byte floats that is 10,000,000 times 1024 times 4, about 41 GB of raw vectors; at 512 it is roughly 20 GB, and at 256 roughly 10 GB. Project to forty million and those become around 164, 82, and 41 GB. On top of the raw vectors, the index structure itself carries overhead. A graph-based index such as HNSW stores neighbour links per vector, and that overhead scales with the number of vectors too. The vectors dominate a large index’s footprint, so dimension, the multiplier on those vectors, is the number with the most leverage.

With the money sized, the next step is measurement. Embed a representative sample at 1024, 512, and 256, build an index for each, and run the same query set with known-relevant passages through all three, scoring recall@k or the downstream answer quality. If 512 holds the quality bar, you have halved the largest line in the storage and memory bill for little or no loss, which on a forty-million-vector index is a large, permanent saving. If 256 also holds, take it. If quality falls off between 1024 and 512, your corpus needs the fidelity and the larger index is justified. The measurement settles it. Assuming bigger is better keeps a storage bill you do not need, and assuming smaller is fine ships worse answers.

Whatever dimension you land on, get the metric and normalisation right once and consistently. Pick cosine similarity unless you have a specific reason not to, leave Titan v2’s normalize flag at its default of true, and build the index with the matching metric. Then keep the query path identical: same model, same dimension, same normalisation, same metric on both the indexing and the querying side. A mismatch here does not throw an error. It retrieves worse passages, and the output still reads like an answer, which is what makes it hard to notice. Re-embedding at a new dimension is a full index rebuild, so plan the cutover rather than trying to change size in place.

Worked example

Take the index at its projected forty million passages. At 1024 dimensions and 4-byte floats the raw vectors are 40,000,000 times 1024 times 4, about 164 GB, before index overhead. The HNSW graph’s neighbour links sit on top of that. Memory is the binding constraint: OpenSearch Service gives half an instance’s RAM to the Java heap and allows k-NN half of what remains, so a node with 32 GiB of RAM holds around 8 GiB of graph. Query latency at the p99 is climbing because larger vectors mean more arithmetic per comparison and a bigger structure to traverse.

The team samples two million passages, embeds them at 1024, 512, and 256, and measures recall@10 against a curated set of a few hundred real questions with hand-checked relevant passages. Recall@10 comes back at 0.94 for 1024, 0.93 for 512, and 0.88 for 256. The application’s floor is 0.90. That settles it: 512 gives up a single point of recall and stays above the bar, while 256 falls below it on this corpus. They re-embed at 512 and rebuild the index. The raw vectors fall to about 82 GB, the working set comes back within a comfortable memory envelope, p99 latency eases, and the monthly storage line roughly halves, for a one-point recall drop the quality bar absorbs. Had the numbers landed differently, with 512 falling to 0.87, the same experiment would have told them to stay at 1024 and instead look at quantisation for savings. Either way the number came from measurement rather than assumption.

Choosing an embedding dimension by corpus and quality Three inputs (corpus size, quality bar, storage and latency budget) feed a measurement step, then a gate that takes the smallest dimension clearing the recall floor, leading to three outcomes (256, 512, 1024) and a closing panel on metric, normalisation and quantisation. Sizing the dimension Inputs Corpus size (now and growth) Quality bar (recall@k) Storage and latency budget Measure each dimension Embed a sample at 1024, 512, 256; score recall on real queries with known good answers. Smallest that clears bar? Take the lowest dimension whose recall stays above the floor; footprint and latency drop with it. 256 Huge, cost-led index; quality still clears bar. 512 Common sweet spot; about half the footprint. 1024 Dense corpus needs the fidelity; baseline. Then lock the invariants and stack savings Match metric (cosine), normalise, rebuild index on change. Same model and dimension across index and query path. Layer quantisation (int8 / binary) as a second, measured saving.

What’s worth remembering

  1. Embedding dimension multiplies against every vector in the index, so on a large corpus it is the dominant term in storage, memory, and per-comparison search cost.
  2. Amazon Titan Text Embeddings v2 supports output dimensions of 1024, 512 and 256, with 1024 the default; Titan G1 is fixed at 1536 and Cohere Embed v3 at 1024, so configurability is a property of the model.
  3. Size the money first with vectors times dimension times bytes-per-value; ten million 1024-d float vectors are about 41 GB of raw vectors before index overhead.
  4. Measure recall on your own data at each candidate dimension against real queries; do not assume bigger is better or smaller is fine.
  5. The distance metric and the model must match what the index expects, and every vector in an index must share the same model and dimension.
  6. Quantisation is a separate saving that stacks on dimension: Titan v2 emits binary embeddings, and a knowledge base takes an embeddingDataType of FLOAT32 or BINARY.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.