PAPER DEEP DIVE
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.
Paper Information
Title: LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
Authors: Xingjun Wang, Gongsheng Li, Qi Fan, Yunlin Mao, Luyan Su, Yingda Chen (ModelScope Team, Alibaba Group)
Link: arXiv:2608.16185
Code Status: No public code repository is provided. The paper page contains no GitHub link. Although the authors are affiliated with Alibaba's ModelScope team, no corresponding implementation was found on ModelScope or GitHub as of writing.
One-Sentence Summary
LENS reformulates in-context search for LLM agents over dynamic document collections as Budgeted Evidence Localization—maintaining query-conditioned beliefs without any pre-built index, using an LLM relevance oracle to iteratively converge toward high-posterior evidence regions, achieving answer quality on par with ReAct while dramatically improving evidence traceability.
Background and Motivation
Large language model agents increasingly need to answer questions over external document collections. But real-world document collections are dynamic—files can be added, updated, or deleted, and these changes occur before a user's query arrives. Traditional retrieval-augmented generation (RAG) approaches materialize the evidence space before the query: through fixed chunking, dense embeddings, persistent sparse indexes, summary trees, or graph-like memory structures.
This pre-materialization strategy is effective when the corpus is stable, because preprocessing costs can be amortized. But in dynamic document settings, it introduces three fundamental tensions. First, setup and update costs must be paid before any query is posed—indexes must be built and maintained even when no one is asking. Second, indexes become stale after document changes—embeddings and chunks no longer match the current document state. Third, fixed chunks commit to an evidence granularity before the query reveals what evidence is needed—but different queries may require a paragraph span, a table entry, a section, or a cross-document chain, whose appropriate granularity depends on the query itself.
The central difficulty is that the evidence space induced by raw documents is latent, variable-boundary, dynamic, and structured. Latent—because answer-bearing evidence exists in the documents but is not known before querying. Variable-boundary—because useful evidence windows are not limited to a fixed set of chunks, making the space finite but combinatorially large. Dynamic—because document updates change the space itself. Structured—because lexical anchors, layout, paths, and historical search signals induce non-uniform priors over likely evidence regions. Treating this space as a fixed finite collection of chunks obscures the actual search problem faced by an LLM agent.
LENS proposes a fundamentally different approach: rather than pre-materializing the evidence space, it forms a low-cost prior to compress the search domain after the query arrives, then uses sequential observations from an LLM relevance oracle to iteratively converge toward high-posterior evidence regions—all under a bounded budget. This makes LENS query-ready immediately after corpus changes, with no index rebuild, while preserving evidence source traceability.
Preliminaries
Understanding LENS requires several key concepts. First is the "latent evidence space"—for a document collection $\mathcal{D}_t$ at time $t$, the set of all possible evidence windows $(d, s, e)$ constitutes $\mathcal{E}_t$, where $d$ is a document and $s$, $e$ are start and end positions. This space is never explicitly enumerated; it is latent—LENS never lists all possible windows.
Second is "query intent"—LENS uses an intent variable $I(q) \in \{\text{lookup}, \text{computation}, \text{comparison}, \text{aggregation}, \text{summarize}\}$ to classify query types. Intent determines data requirements: lookup queries need $K=1$ atomic fact, while comparison and aggregation queries require $K > 1$ facts that may reside in different documents.
Third is the "oracle"—LENS treats the LLM as a costly relevance oracle. Each call to the LLM to judge a candidate evidence region's relevance to the query consumes tokens, latency, and cost. The search method must therefore balance immediate relevance, information gain, source traceability, and budget.
Finally, "belief update"—LENS maintains a posterior belief $P(Z_j^* \mid f_j, q, \mathcal{H}_t)$ for each atomic fact $f_j$, where $\mathcal{H}_t$ is the observation history. Each oracle observation $o_i$ simultaneously updates all facts' beliefs—this is the key advantage of the per-fact decomposition.
Method Details
Layer 1: Low-Cost Prior
LENS's first layer compresses the search space by fusing five families of low-cost signals available without reading the corpus through an LLM: lexical anchors, document-path structure, compiled document summaries when available, historical source-grounded evidence from prior successful searches, and lightweight corpus scans. Rather than treating these signals as independent retrieval modules, LENS uses them as approximations to a prior:
$$\pi_{\text{prior}}(z \mid q, \mathcal{D}_t) \approx \sum_{k \in \mathcal{K}_0} w_k \, \pi_k(z \mid q, \mathcal{D}_t) \tag{6}$$
where $\mathcal{K}_0$ is the family of low-cost proposal signals. This joint prior factors, by the chain rule, into a marginal–conditional pair that separates document selection from within-document localization:
$$\pi_{\text{prior}}(z \mid q, \mathcal{D}_t) = \pi_{\text{file}}(d_z \mid q, \mathcal{D}_t) \, \pi_{\text{pos}}(s_z, e_z \mid d_z, q) \tag{7}$$
where $d_z$ is the document associated with region $z$. This decomposition matters because strong file-level evidence does not automatically imply precise within-document localization. Proposition 1 proves that before any iterative oracle budget is consumed, LENS reduces the effective search domain from the full raw-document collection to a query-conditioned subspace $\mathcal{C}_{\text{search}}$. Corollary 1 gives the concrete magnitude: with file-admission width $m=10$ and comparable article lengths, the reduction ratio is approximately $\mathcal{C}_{\text{search}} / |\mathcal{E}_t| \approx 1.8 \times 10^{-3}$ for $D_{125}$ down to $4.7 \times 10^{-4}$ for $D_{500}$—roughly three orders of magnitude.
Layer 2: Budget-Constrained Sequential Inference
After prior formation, LENS enters the budgeted exploration loop with four steps cycling: propose a candidate evidence region $z_t$, query the LLM relevance oracle to obtain observation $o_t$, update beliefs over evidence regions, and adapt proposal weights and coverage estimates. The observation history $\mathcal{H}_t = \{(z_i, o_i)\}_{i=1}^t$ maintains one belief per fact and updates all of them from the shared history:
$$P(Z_j^* \mid f_j, q, \mathcal{H}_t) \propto \prod_{i=1}^t P(o_i \mid Z_j^*, z_i, f_j, q) \cdot \pi_{\text{prior}}(Z_j^* \mid f_j, q, \mathcal{D}_t) \tag{8}$$
A single oracle call therefore serves all facts at once—this is why the per-fact formulation does not multiply the oracle budget by $K$.
The next candidate region should balance exploitation and exploration. An ideal information-directed objective selects:
$$z_{t+1} = \arg\min_{z \in \mathcal{C}_{\text{search}}} \Psi_t(z), \quad \Psi_t(z) = \frac{[\Delta_t(z)]^2}{\mathbb{I}(\mathbf{Z}^*; O_z \mid q, \mathcal{H}_t)} \tag{9}$$
where $\Delta_t(z)$ is the expected immediate relevance gap and the denominator is the expected information gain about the outstanding targets $\mathbf{Z}^*$. Minimizing this information ratio is principled: it targets the trade-off between immediate relevance and long-run information gain that underlies regret-optimal sequential selection. Since exact computation is intractable over raw documents, LENS approximates this criterion with complementary proposal families:
$$\pi_t(z) = \lambda_{\text{lex}}^{(t)} \pi_{\text{lex}}(z) + \lambda_{\text{local}}^{(t)} \pi_{\text{local}}(z) + \lambda_{\text{global}}^{(t)} \pi_{\text{global}}(z) \tag{10}$$
Lexical proposals exploit anchors, local proposals refine around high-belief regions, and global proposals guard against semantic omissions. Treating each proposal family as an arm, LENS adapts the mixture weights $\lambda^{(t)}$ online from observed oracle utility, closing the propose–observe–update cycle without enumerating the full latent space.
Figure 1: Overall LENS framework. LENS forms a query-conditioned prior over candidate evidence regions, runs a budget-constrained propose–observe–update loop, and consolidates selected regions into a compact source-grounded evidence set for answer synthesis. It never pre-materializes a persistent index over the raw-document collection.
Budget-Aware Stopping
The exploration loop should not continue merely because more context can be read. LENS exits when either the remaining budget is insufficient or every requirement is localized with sufficiently concentrated belief. Following fixed-confidence best-arm identification ideas, a conceptual per-fact stopping statistic is:
$$\text{GLR}_t^{(j)} = \min_{z \neq \hat{Z}_{j,t}} \sum_{i \leq t} \log \frac{P(o_i \mid \hat{Z}_{j,t}, f_j, q)}{P(o_i \mid z, f_j, q)} \tag{11}$$
where $\hat{Z}_{j,t}$ is the current highest-belief region for fact $f_j$. LENS stops when $\min_{j \leq K} \text{GLR}_t^{(j)}$ exceeds an intent-modulated threshold $\beta(t, \delta) \, \gamma(I)$—that is, when the weakest requirement is resolved. The intent factor $\gamma(I)$ tightens the criterion for computation and comparison intents and relaxes it for lookup. A lookup query has a single requirement and can often stop with one compact region, whereas comparison and computation queries cannot stop until each of their $K$ facts has its own confirmed window—exactly the coverage condition on $\mathcal{D}_{\text{req}}(q, I)$.
Evidence Consolidation and Answer Synthesis
Once the loop stops, selected regions enter the consolidation-and-synthesis stage. Consolidation merges the per-fact windows $\{\hat{Z}_j\}_{j \leq K}$, removes redundant or overlapping regions, expands boundaries when necessary for interpretability, preserves source traces, yielding a compact source-grounded evidence set $E^*$. Answer synthesis then operates on $E^*$: for computation and comparison queries, LENS separates extraction of atomic facts from answer synthesis; for lookup-style queries, a single-stage synthesis may be sufficient. When synthesis cannot satisfy $\mathcal{D}_{\text{req}}(q, I)$ and budget remains, LENS triggers the self-correction path: a bounded step that relaxes the stopping threshold and re-enters the exploration loop with an expanded candidate set before re-synthesizing. The final output is the pair $(E^*, a)$—an answer grounded in explicit evidence regions rather than only in retrieved text snippets.
graph TD
A["Query q + Dynamic Corpus D_t"] --> B["Layer 1: Low-Cost Prior"]
B --> B1["Lexical Anchors"]
B --> B2["Document Path Structure"]
B --> B3["Compiled Summaries"]
B --> B4["Historical Evidence"]
B --> B5["Lightweight Scans"]
B1 & B2 & B3 & B4 & B5 --> C["Candidate Subspace C_search"]
C --> D["Layer 2: Budget-Constrained Loop"]
D --> D1["Propose Candidate z_t"]
D1 --> D2["LLM Relevance Oracle"]
D2 --> D3["Update Per-Fact Beliefs"]
D3 --> D4["Adapt Proposal Weights"]
D4 --> E{"Budget Left? Requirements Met?"}
E -->|Yes| D1
E -->|No| F["Evidence Consolidation"]
F --> F1["Merge & Deduplicate"]
F1 --> F2["Source-Grounded Evidence Set E*"]
F2 --> G["Answer Synthesis"]
G --> G1{"Requirements Satisfied?"}
G1 -->|No, Budget Left| H["Self-Correction: Relax Threshold
Re-enter Loop"]
H --> D
G1 -->|Yes| I["Output (E*, a)"]
Figure 2: LENS algorithm flowchart. From low-cost prior formation through the budget-constrained exploration loop to evidence consolidation and answer synthesis, with the self-correction path re-entering the loop when requirements are unmet.
Algorithm Summary and Theoretical Properties
Algorithm 1 summarizes the inference loop. Proposition 2 proves bounded oracle complexity: for a loop budget of $L$ exploration rounds, the number of LLM oracle interactions performed by LENS is bounded by $c_0 + c_1 L$, where $c_0 = 4$ and $c_1 = 2$ in the current configuration—independent of the number of requirements $K$ (because one oracle observation updates all per-fact beliefs) and independent of $|\mathcal{E}_t|$ (the number of latent evidence windows induced by the documents). At the actual configuration $L=3$, the theoretical upper bound is 10 oracle interactions per question; measured averages are only 4.00 (G125/G250) and 4.04 (G500), with zero budget-exceeded records. The slack is expected—most questions terminate through a short-circuit (coverage check reports completion, or an intent-gated direct-analysis path resolves the query before the loop is exhausted).
The oracle interaction count by stage: S1 anchor extraction (1 request), S2 requirement decomposition (1), S3 exploration rounds (2 per round—proposal-and-observation + coverage), S4 answer synthesis (1), S5 answer-span calibration (≤1), totaling $1 + 1 + 2L + 1 + 1 = 4 + 2L$. The key property: this bound does not grow with corpus size or requirement count—the fundamental guarantee of LENS's scalability in large-scale dynamic document settings.
Experimental Results
Setup
LENS is evaluated on HotpotQA fullwiki—a multi-hop question-answering benchmark where each question requires reasoning over two or more Wikipedia articles. Two conditions are reported: controlled evaluation ($D_n$, drawn from the validation split of 7,405 questions with stratified sampling, seed 42) and open-domain fullwiki (the complete raw Wikipedia dump stored as 15,517 JSON shards without preprocessing, chunking, or indexing).
Five systems are compared: LENS (full algorithm), ReAct Search (strong iterative baseline using ReAct-style tool-use reasoning without LENS's structured prior or budget-constrained belief updates), Hybrid-RAG (BM25 + dense embedding retrieval with pre-materialized index), BM25-RAG (sparse-retrieval baseline with pre-built BM25 index), and Closed-Book (no-retrieval reference estimating model parameter contribution alone). All systems share one chat backend (Qwen3.7, a 35B-parameter MoE model with 3B active parameters) under a 300-second per-question wall-clock cap. LENS runs its DEEP configuration with a 128K query-time token budget and at most 10 candidate files admitted to evidence extraction.
| System ($D_{500}$) | Query-Ready | Rebuild | Index(s) | Storage |
|---|---|---|---|---|
| LENS | Yes | No | 0.0 | 0 |
| ReAct | Yes | No | 0.0 | 0 |
| BM25-RAG | No | Yes | 4.0 | 10.0MB |
| Hybrid-RAG | No | Yes | 5.7 | 54.9MB |
Table 1: $D_{500}$ query readiness, rebuild requirement, index build time, and index storage. LENS and ReAct require no pre-built index.
Main Results: Controlled Evaluation
On the $D_{500}$ controlled evaluation with 500 questions, ReAct Search attains the highest answer score (65.2% EM, 78.9% F1), while LENS remains close on answer quality (62.4% EM, 76.9% F1) and provides substantially stronger evidence localization. LENS achieves 84.8% evidence recall and 96.8% grounded answers, compared with 50.4% and 71.8% for ReAct. This separates two evaluation dimensions: ReAct more frequently produces the exact answer string, whereas LENS more reliably localizes and traces the supporting evidence.
| System | EM | F1 | Ev.Rec | Ground |
|---|---|---|---|---|
| ReAct Search | 65.2 | 78.9 | 50.4 | 71.8 |
| LENS | 62.4 | 76.9 | 84.8 | 96.8 |
| Hybrid-RAG | 38.4 | 51.1 | 80.8 | 92.2 |
| BM25-RAG | 28.8 | 42.3 | 71.8 | 95.8 |
| Closed-Book | 35.2 | 47.0 | 0.0 | 0.0 |
Table 2: Controlled evaluation on $D_{500}$ (%, n=500). EM/F1 are official answer metrics; Ev.Rec is supporting-fact document recall; Ground is the percentage of answers traceable to retrieved evidence.
Relative to Closed-Book, LENS gains 27.2 percentage points EM and ReAct gains 30.0. Paired McNemar tests show no significant LENS–ReAct answer-quality difference on $D_{500}$ (p=0.1143) or on fullwiki dev-150 (p=1.0000).
Open-Domain Fullwiki Results
On the complete raw Wikipedia corpus (15,517 shards), LENS and ReAct are effectively tied on official answer quality (43.3% vs. 42.7% EM), while LENS grounds a larger share of answers in retrieved evidence (84.0% vs. 70.7%). Closed-Book reaches 38.7% EM, so fullwiki retrieval gains are modest but positive: +4.6 pp for LENS and +4.0 pp for ReAct.
| System | EM | F1 | Ev.Rec | Ground |
|---|---|---|---|---|
| LENS | 43.3 | 57.5 | 45.0 | 84.0 |
| ReAct Search | 42.7 | 57.3 | 50.4 | 70.7 |
| Closed-Book | 38.7 | 49.7 | 0.0 | 0.0 |
Table 3: Open-domain fullwiki results (%, n=150, fixed sample IDs). LENS and ReAct search the raw Wikipedia dump with zero indexing.
Evidence Recall Leadership
Across all controlled evaluation scales, LENS is consistently the top evidence-localization system. Its lead over ReAct is 43.0 pp on $D_{125}$, 38.0 pp on $D_{250}$, and 34.4 pp on $D_{500}$. Hybrid-RAG also retrieves substantial evidence, but its answer quality remains far below LENS and ReAct, indicating that locating candidate evidence and synthesizing the exact multi-hop answer are separable failure modes.
Corpus Staleness and Lifecycle Robustness
The stale-index arm measures performance when a $D_{125}$-built index is reused after the corpus expands to $D_{250}$. BM25-RAG and Hybrid-RAG are index-dependent and therefore lose most of their ability to answer newly added questions: EM drops by 28.0 and 28.8 pp, evidence recall drops by 70.1 and 69.6 pp. ReAct and LENS are index-free and can query the expanded corpus immediately—EM changes are small (+0.8 and -2.4 pp), and LENS retains nearly all supporting-fact recall (84.7% to 83.9%).
Ablation Study
The ablation is conducted on the fullwiki fixed 150-question subset. Removing sequential exploration produces the clearest degradation: EM falls from 43.3% to 38.0%, F1 from 57.5% to 50.9%, and evidence recall from 45.0% to 31.0%. Removing the multi-signal prior does not reduce EM on this subset, but it lowers F1 and changes the cost profile. This confirms that the sequential exploration loop is the primary contributor to LENS's answer quality and evidence localization.
Limitations
First, the evaluation scope is limited. HotpotQA emphasizes lookup and comparison queries over encyclopedic text—richer document layouts, aggregation intents, table evidence, and warm-reuse behavior remain future work. The authors explicitly acknowledge this: "HotpotQA emphasizes lookup and comparison over encyclopedic text, so richer document layouts, aggregation intents, table evidence, and warm-reuse behavior remain future work."
Second, the theoretical analysis provides structural statements rather than tight guarantees. Proposition 1 is a cost statement (not a sufficiency claim); Proposition 2's oracle bound is a control-flow analysis (not a mean predictor). Stricter probabilistic oracle models, position-level priors, and resampling analyses remain future work. The per-fact update's exactness depends on two assumptions (A1 prior independence, A2 conditional independence of observations), which can fail when requirements are logically coupled or when the LLM's per-requirement verdicts are correlated—under such conditions, the per-fact update should be treated as a tractable approximation.
Third, prior compression has a price. Corollary 1 of Remark 1 notes that all posterior mass is confined to $\mathcal{C}_{\text{search}}$: if $Z_j^* \notin \mathcal{C}_{\text{search}}$ for some requirement $f_j$, then $P(Z_j^* \mid f_j, q, \mathcal{H}_t) = 0$ for all $t$, and no amount of round budget can recover it. File-level compression converts a $\sim N/m$ saving into a recall ceiling determined by the quality of $\pi_{\text{file}}$. The bounded self-correction path partially lifts this ceiling, but this is an inherent architectural limitation.
Fourth, LENS consumes more tokens than ReAct (16.5K vs. 11.8K tokens/query on $D_{500}$), but gains 34.4 pp evidence recall and 25.0 pp grounding. In latency-sensitive scenarios, this trade-off requires re-evaluation.
Conclusion and Outlook
LENS's core contribution is reformulating in-context search for LLM agents over dynamic document collections from "one-shot top-k retrieval" to "budgeted sequential evidence localization." This reformulation has three key implications: the evidence space is latent rather than pre-materialized, evidence granularity is query-conditioned rather than preset, and the search process is bounded sequential inference rather than single-shot lookup.
On controlled corpora, LENS is the strongest evidence-localization system: on $D_{500}$ it trails ReAct by 2.8 pp EM but leads by 34.4 pp evidence recall and 25.0 pp grounding. On fullwiki dev-150, the two systems are effectively tied in official EM/F1, while LENS retains stronger grounding. On the stale-index arm, index-dependent systems lose 28.0–28.8 pp EM and nearly all evidence recall, while index-free systems remain query-ready.
LENS is positioned not as an EM-dominant answer generator, but as an index-free evidence localization method that makes answers more traceable to current raw sources. EM as a narrow string-match metric penalizes acceptable paraphrases and aliases, and can reward a correct answer produced from model memory rather than from retrieved evidence—this is exactly what the Closed-Book reference baseline makes visible.
Future directions include: extending to richer document layouts and table evidence, aggregation-intent evaluation, warm-reuse behavior research, stricter probabilistic oracle models, and position-level priors with resampling analyses. These extensions will validate the LENS framework's applicability across a broader range of dynamic document scenarios.
Golden Quote
"LENS is not an EM-dominant answer generator, but an index-free evidence localization method that makes answers more traceable to current raw sources." — A correct answer without traceable evidence is insufficient for settings that demand auditability. LENS chooses a different path: sacrifice a few points of exact-match, gain a massive improvement in evidence traceability.
SOURCE LINKS


