TL;DR
- Mixedbread Search is an API service that builds searchable multimodal indexes from files such as PDFs, images, audio, and video.
- I was curious about how the product might address two problems.
- Based on several assumptions, I wrote down how I would approach them through a few small experiments (customer-specific retrieval signals and web retrieval dependency).
- Mixedbread may already have solved these problems. Still, thinking through them was fun!
1. Introduction
Recently, I have been interested in agentic search and exploring the topic from different angles. Along the way, I came across a YouTube talk by Mixedbread AI’s Hanna Lichtenberg and Aamir Shakir titled How We Taught Agents to Use Good Retrieval.1
I had recently been thinking about similar problems myself, so the insights from the talk immediately caught my attention. I was already familiar with the Mixedbread name from having used one of its embedding models, and the talk made me curious to learn more about the company. I started by looking through its products and blog posts.
Two questions came to mind while I was exploring the product:
- How can retrieval quality be measured and improved for each customer?
- How should a dependency on web retrieval be managed?
The rest of this post discusses both questions in more detail. Mixedbread may already have solutions to these problems, but my goal is simply to organize how I would approach them.
2. Understanding Mixedbread Search
2.1 Expected Product Objective
Mixedbread Search is an API-based service that creates a searchable Store from uploaded PDFs, images, documents, code, audio, and video.2 It provides natural-language Search, Agentic Search, and Question Answering. Customers do not need to build their own parsing, chunking, embedding, indexing, and retrieval pipelines.
In other words, the product ultimately aims to make retrieval infrastructure something customers do not have to think about.
Based on this understanding, I expect the product objectives to be:
- Accurately retrieve the evidence a customer needs from its heterogeneous private documents.
- Manage latency, storage, and inference costs while maintaining high retrieval quality.
2.2 Mixedbread’s Technical Strengths
Retrieval Models. Mixedbread has developed Wholembed v3, a unified omnimodal, multilingual late-interaction retrieval model.3 It handles text, images, audio, and video with a single retrieval model.
Reranking. Mixedbread also provides its own listwise reranker.4 It ranks a candidate set as a whole and supports instruction-based ranking criteria such as relevance, freshness, and source preference.
Serving and Retrieval Infrastructure. Silo is Mixedbread’s retrieval engine for serving multimodal late-interaction retrieval at billion scale.5 maxsim-cpu optimizes MaxSim operations for CPU serving,6 while asymmetric quantization reduces storage and retrieval costs.7
Agentic Search. Mixedbread’s Agentic Search supports query decomposition, parallel sub-query search, wide search, and multi-round retrieval.8 Mixedbread also uses supervised fine-tuning and reinforcement learning to teach its search agent to use retrieval tools more effectively.1
Benchmarks. According to Mixedbread’s published evaluations, its system ranks first on BrowseComp-Plus. It also performs strongly on multimodal and enterprise agentic retrieval benchmarks such as MADQA and OfficeQA-Pro.9
2.3 Inferred High-Level Architecture
The following architecture is my inference based on Mixedbread’s public materials.
Indexing Flow
Customer files
↓
File-type-specific parsing
(OCR, layout analysis, transcription, and media processing)
↓
Chunking and metadata generation
↓
Multimodal late-interaction encoding
(Wholembed v3)
↓
Customer-specific Store or index
(S3-based)
After a file is uploaded, asynchronous workers likely perform parsing, chunking, embedding, and indexing. I also expect the system to build an approximate nearest neighbor (ANN) index.
According to Mixedbread’s post on billion-scale multimodal late-interaction retrieval, S3 serves as the source of truth for index data. At query time, the required data is loaded onto local NVMe SSDs and cached in memory.5
Logical isolation between Stores clearly exists, but the physical sharding and tenant-isolation strategies are not publicly documented.
Search Flow
Search request
├─ Standard Search
│ └─ Customer-specific index
│ + external web API when requested
├─ Agentic Search
│ └─ Agentic loop using Standard Search
│ + query decomposition & reranking by the agent
└─ Question Answering
└─ Standard Search
+ answer & citation generation
Standard Search retrieves candidates from an internal index, with optional query rewriting and reranking.
Agentic Search performs multiple rounds of query planning and retrieval. Question Answering uses retrieved results as context to generate an answer with citations.
How the Web Store10 obtains its search candidates is not publicly documented. It could rely on its own selectively maintained index, external search providers, or a combination of both. One possibility is that it uses an external search provider with a strict timeout to stay within its latency budget. In the later proposal, I consider the case where an external provider is involved.
3. Two Questions
Mixedbread’s public benchmarks demonstrate that its retrieval stack performs well across general, multimodal, and enterprise workloads. However, operating the product in production raises two additional questions.
3.1 How Can Retrieval Quality Be Measured for Each Customer?
Each customer has different documents, vocabulary, query distribution, and definition of relevance. Strong performance on public benchmarks therefore does not necessarily guarantee that every customer-specific Store continues to work well on its actual workload.
This creates a measurement problem. Mixedbread may observe the search query and returned documents, but it may not directly observe whether the final user was satisfied or whether the retrieved evidence led to a correct answer.
The first question is therefore how to generate useful relevance signals (or reliable proxies) from customer-specific traffic and documents.
Such signals could support:
- Customer-specific evaluation sets
- Retrieval regression monitoring
- Failure discovery and analysis
- Selection of indexing and retrieval configurations
- Training or calibration of query rewriting and reranking components
The broader goal would be to create a feedback loop in which each Store can be evaluated and gradually improved as its documents and query distribution evolve.
3.2 How Should a Dependency on Web Retrieval Be Managed?
Mixedbread also provides a Web Store that can be searched independently or together with private Stores. Its candidate acquisition mechanism is not public, but one possible implementation is to rely partly on an external web search provider.
If such a dependency exists, several challenges arise:
- Changes to the external provider may alter which documents are available to Mixedbread’s downstream ranking system, causing retrieval quality to change even when Mixedbread’s own components remain unchanged.
- Queries that work well for private documents may not be appropriate for the web, where public terminology, freshness constraints, entity disambiguation, and source preferences may matter more.
- Repeated calls to an external provider may increase latency and API costs, particularly during multi-round agentic search.
These lead to three practical questions:
- How can Mixedbread detect whether a quality regression originates from the external provider or its own downstream ranking?
- Should private and web retrieval use different query-generation policies?
- How can repeated web retrieval be cached without returning stale evidence?
The following sections explore each of these questions through a concrete evaluation or optimization experiment.
4. Building Customer-Specific Retrieval Signals
The core challenge is obtaining a useful relevance signal for each customer’s actual workload.
The ideal signal would tell us whether the retrieved documents satisfied the user’s information need. However, Mixedbread may only observe the search query and returned documents, while the final answer and user behavior remain inside the customer’s application.
There are several possible ways to address this.
4.1 Explicit Feedback
The most reliable option is to let customers send retrieval feedback directly through an API.
For example, customers could report that:
- A result was relevant or irrelevant.
- An expected document was missing.
- A search succeeded or failed.
- A particular result was used in the final answer.
One possible way to encourage customers to provide this feedback would be to offer lower API costs, evaluation reports, or customer-specific search optimization.
This is primarily a product and incentive-design problem rather than a research problem, so I will not focus on it further here.
4.2 Implicit Signals from Query Behavior
When explicit feedback is unavailable, query sequences may provide weak signals of dissatisfaction.
Examples include:
- Repeating the same query while changing
top_kor another search parameter. - Rephrasing a query shortly after the initial search.
- Broadening or narrowing the query.
- Switching from Standard Search to Agentic Search after an initial attempt.
None of these behaviors proves that the first result was poor. However, they can help identify query-result pairs that are more likely to contain retrieval failures and are therefore worth further evaluation.
4.3 Estimating Relevance
Once potentially problematic queries have been identified, relevance labels can be created in several ways.
Human Labeling. For large customers, it may be worthwhile to label a sample of important queries manually. Human judgments provide the most reliable signal, especially in specialized domains. However, they are expensive, and external annotators may not understand customer-specific terminology or relevance criteria.
LLM-as-a-Judge. An LLM judge can evaluate whether a retrieved document is relevant to a query. This makes it possible to label a large number of query-document pairs at relatively low cost. The judge itself could be designed and calibrated using the process I described in LLM-as-a-Judge in Practice.
However, evaluating only the documents returned by the current system creates selection bias. The judge may determine whether retrieved documents are relevant, but it cannot identify relevant documents that the system never retrieved.
A Stronger Offline Search Agent. To discover documents that the production system may have missed, a more expensive search process could run offline.
Unlike the serving system, the offline process would not need to meet the same latency or cost constraints. It could use:
- More query expansions.
- A larger candidate pool.
- Multiple retrieval configurations.
- Exact keyword and semantic search.
- Full-document inspection.
- More search rounds.
The union of these results could be judged by a stronger model or verified by humans to construct a more complete set of relevant documents.
This would not create perfect ground truth, but it could provide a stronger reference than the production system alone.
4.4 From Signals to a Feedback Loop
The resulting explicit or proxy signals could be used to build a customer-specific evaluation set.
Customer queries
↓
Explicit feedback, behavioral signals,
LLM judgments, or stronger offline search
↓
Customer-specific evaluation set
↓
Failure discovery and regression monitoring
↓
Index and retrieval improvements
↓
Evaluation on newly collected queries
The evaluation set should be updated periodically as the customer’s documents and query distribution change.
It could then support:
- Detecting regressions after model or index updates.
- Comparing chunking and metadata strategies.
- Selecting retrieval depth and reranking configurations.
- Improving query rewriting.
- Mining hard examples for retriever or reranker training.
The important point is not only to measure customer-specific retrieval quality, but to use the measurement to keep each Store from becoming stale as its workload evolves.
4.5 A Minimal Experiment
A small initial experiment could focus on a few customers with sufficiently large query volumes.
- Sample real search queries from each customer.
- Identify likely failure cases using query-sequence signals.
- Generate relevance labels using an LLM judge and a stronger offline search process.
- Verify a small subset with human labels.
- Use the resulting evaluation set to identify one recurring failure pattern.
- Improve one component, such as query rewriting, metadata contextualization, or reranking.
- Evaluate the change on held-out customer queries.
The primary metric could be customer-specific Recall@K or nDCG@K. Latency and cost should be monitored as constraints to ensure that the improvement remains practical.
5. Managing a Dependency on Web Retrieval
If Mixedbread’s Web Store feature relies on an external search API, three potential problems arise.
5.1 Robustness to Upstream Search Changes
If the Web Store relies on an external search provider, Mixedbread does not fully control which documents enter its ranking pipeline. A provider update could change the candidate set even when Mixedbread’s own system remains unchanged.
When final retrieval quality drops, there are two possible causes:
- The provider no longer retrieves the relevant documents.
- The relevant documents are still retrieved, but the downstream ranking system ranks them poorly.
Building a Web Retrieval Evaluation Set
A pooled relevance set could be built from a representative sample of real Web Store queries.
For each query, candidates could be collected from the current provider, historical snapshots, an alternative provider if available, and several query rewrites. The union would then be deduplicated and labeled for relevance using an LLM judge, with a subset manually verified.
Pooling candidates from multiple sources is important because evaluating only the current provider’s results cannot reveal relevant documents that the provider failed to retrieve entirely.
Complementary Metrics
Two metrics can then separate upstream retrieval quality from downstream ranking quality:
- Candidate Recall@K: whether the external provider includes the relevant documents in its candidate set.
- nDCG@K: whether the downstream ranking system places those relevant documents near the top of the final results.
They are complementary because nDCG alone cannot tell whether a relevant document was ranked poorly or never entered the candidate set.
Their combination also makes regressions easier to diagnose:
Candidate Recall ↓
→ Upstream provider is losing relevant documents.
Candidate Recall stable + nDCG ↓
→ Relevant documents are still available,
but downstream ranking has degraded.
Candidate Recall stable + nDCG stable
→ The provider change has not materially affected retrieval quality.
This distinction also suggests different responses: upstream regressions may require query rewrites, deeper retrieval, or a fallback provider, while downstream regressions may call for reranker recalibration or retraining.
A Minimal Experiment
A small initial experiment could use a few hundred representative Web Store queries.
- Sample queries from real Web Store traffic.
- Retrieve results from the current provider.
- Add candidates from an alternative source and several query rewrites.
- Merge, deduplicate, and label the resulting candidate pool.
- Save the current provider output as the baseline snapshot.
- Re-run the same queries periodically or after a provider update.
- Compare Candidate Recall@K and final nDCG@K against the baseline.
This would reveal not only whether web retrieval quality changed, but also whether the regression originated upstream or in Mixedbread’s downstream ranking.
5.2 Different Query-Generation Policies for Private and Web Retrieval
Queries that work well for private documents may not work equally well on the web because the two sources contain different kinds of information.
Consider the question:
Does our parental leave policy comply with current California law?
For the private Store, useful queries may focus on the company’s own terminology, such as parental leave policy or employee leave handbook.
For the Web Store, the goal is different: retrieving the current external regulation. A query such as California CFRA parental leave 2026, possibly with a preference for official government sources, may be more appropriate.
This suggests that Agentic Search may benefit from generating source-specific queries rather than applying the same rewritten query to both sources.
A Minimal Experiment
A simple experiment could compare two policies on hybrid private-and-web queries:
- Shared query policy: use the same generated query for both private and web retrieval.
- Source-specific policy: generate separate queries for the private Store and the Web Store.
The comparison could measure:
- Evidence Recall@K
- Number of search rounds
- Web search calls
- End-to-end latency
If source-specific generation retrieves the required evidence more reliably or with fewer search rounds and web calls, it would suggest that private and web retrieval should use different query-generation policies.
5.3 Freshness-Aware Caching for Web Retrieval
Repeated calls to an external web provider can add latency and API cost, especially during multi-round Agentic Search. Caching can reduce this overhead, but web results have different freshness requirements.
For example, a query about a stable historical fact may remain reusable for a long time, while a query about current news or prices may become stale within minutes.
A practical approach is to make cache lifetime depend on the expected freshness of the query:
Web query
↓
Freshness estimation
↓
Static → long TTL
Slow-changing → medium TTL
Time-sensitive → short TTL
Real-time → bypass cache
Caching could be applied at multiple levels, such as search results for repeated queries and fetched or parsed documents for repeated URLs.
A Minimal Experiment
Using historical Web Store traffic, compare:
- No cache
- Exact-query cache
- Semantic-query cache
- Freshness-aware semantic cache
Measure:
- External API call reduction
- p95 latency
- Retrieval quality
- Stale-result rate
The goal is to reduce external calls and latency without materially degrading the freshness or relevance of the retrieved evidence.
6. Conclusion
Mixedbread has strong models of its own, the infrastructure to serve them efficiently, and a well-built product that brings those capabilities together.
I thought about what signals this product would need to keep improving over time and explored ways to obtain those signals, even if only through imperfect proxies.
Of course, Mixedbread may already have solved all of these problems, and my assumptions about its system may be completely wrong.
If you have read this far, something in this post probably caught your interest. If you would like to know more about me, feel free to connect with me on LinkedIn.
References
-
Hanna Lichtenberg and Aamir Shakir, How We Taught Agents to Use Good Retrieval. ↩ ↩2
-
Mixedbread, Wholembed v3. ↩
-
Mixedbread, Listwise Reranking. ↩
-
Mixedbread, Multimodal Late-Interaction Retrieval at Billion Scale. ↩ ↩2
-
Mixedbread, maxsim-cpu. ↩
-
Mixedbread, Asymmetric Quantization. ↩
-
Mixedbread, Agentic Search. ↩
-
Mixedbread, Evaluations. ↩