
Building Uber's Legal Redlining Agent: Four Iterations to a Lawyer-Trusted AI
Uber's Legal Redlining Agent lives inside Microsoft Word: four iterations (RAG, lawyer feedback, tone, agentic modifications) cut average review time over 20% with 91% decision accuracy.
Few domains demand as much precision and care as legal work, which makes it one of the hardest possible tests for putting an AI agent into daily professional practice. Beginning in early 2024 and continuing through 2025, before today's agent harnesses were widely available, Uber built the first generation of its Legal Redlining Agent (LRA). The system helps Uber Legal teams handle high-volume contract negotiations without compromising trust or compliance, and without replacing a lawyer's judgment.
The most durable lesson from that effort was not about the model or the architecture. It was about bringing an agent successfully into lawyers' day-to-day work. Uber started from the business problem: a steady stream of contract negotiations involving substantial repeatable redline work. It met lawyers where they already work, inside Microsoft Word, so LRA could flag risky redlines and suggest edits within the existing workflow. It then iterated closely with lawyers on what the agent should do, piloting it with the Legal team in 2025 and folding their feedback back into the product.
This article follows what Uber learned bringing that first generation into practice: framing the problem, earning lawyers' trust, and improving the agent through their feedback. It closes with how the first-generation system worked, and how the team would approach the problem today as agent technology has evolved.
On March 9, 2026, LRA was part of Uber Legal's winning submission for Most Innovative Legal Department of the Year at the ALM Legalweek Leaders in Tech Law Awards.
The problem: high stakes, high friction
At Uber, thousands of contracts are negotiated every year. Behind each one is a legal team reviewing and negotiating the text. The process can be time-consuming, repetitive, and often a bottleneck for deals, onboarding, and launches. It is high-stakes, detail-oriented work where precision matters and delays ripple across sales, onboarding, and launches.
Yet buried inside this high-stakes process was something surprisingly structured: similar clauses, repeatable redlines across clients, and standing legal reasoning patterns. In other words, there was a clear system.
Because the legal organization's work is mission-critical to Uber, the team saw an opportunity to automate the repeatable parts and free up lawyers' time for complex judgment calls, all without sacrificing quality or consistency.
The solution: an AI-powered redlining agent
The Legal Redlining Agent is an AI-powered assistant integrated directly into Microsoft Word as an add-in, where lawyers already work. It reviews client edits, understands intent, recommends policy-aligned responses, generates structured comments, flags risks, and continuously learns from feedback. It combines six capabilities:
- Document review to identify redlines, distinguishing edits made by clients from edits made by Uber lawyers.
- Intent detection to interpret what a proposed change is actually trying to do.
- Policy-aligned recommendations to accept, reject, or modify clauses.
- Comment generation to explain Uber's legal position back to the counterparty.
- Risk flagging to highlight clauses that need deeper review by a lawyer.
- A self-learning feedback loop to improve future recommendations.
Since launching with the Legal team, Uber has seen an over 20% reduction in average contract review time and 91% accuracy in AI-generated decisions. Feedback from lawyers includes: "It saved me real time; especially with comment suggestions."
What is the Legal Redlining Agent?
The add-in analyzes a Microsoft Word document and identifies any redlines from external parties. Using AI, it processes these changes, attempting to understand the client's intent and the legal implications, and compares the proposed changes against the team's legal policies and guidelines.
Based on this analysis, the add-in generates a series of suggestions: recommended actions, proposed comments back to the client that justify these recommendations with legal reasoning, and an assessment of the risk level associated with each change. Uber lawyers then review each suggestion, deciding whether to accept, reject, or modify it before final incorporation into the client response.
Four iterations toward a lawyer-trusted agent
The system described above did not emerge fully formed. Getting there required four major iterations, each of which taught hard lessons about what legal AI actually requires. Along the way, the team used Uber's GenAI Gateway, which allowed them to focus entirely on refining product logic while working closely with subject matter experts.
| Iteration | Focus | Hard lesson |
|---|---|---|
| 1. RAG | Playbook retrieval plus LLM adjudication | Semantic similarity, tone and generalization all failed on real negotiations |
| 2. Feedback | Self-learning loop from lawyer decisions | Granular feedback (expected action, comments, final text) is what made the agent accurate |
| 3. Tone | Dedicated tone-modulation pass | Prompts that shape output voice must be owned by the subject matter experts |
| 4. Agentic modifications | Drafting counter-proposals, rules database | Drafting the compromise language is the real time-saver, and hard policies need deterministic guardrails |
Iteration 1: RAG
The initial approach focused on a foundational RAG (Retrieval Augmented Generation) system. The Legal team provided existing playbooks and examples of negotiation turns. The hypothesis was that by ingesting these playbooks to identify semantically relevant sections, and then feeding those sections along with proposed counterparty changes to an LLM, the system would generate sound decisions.
Three problems appeared quickly:
- Inaccurate semantic similarity. The nature of the "key" being embedded differed significantly from the content being retrieved, producing imprecise similarity searches and unreliable application decisions.
- Inappropriate tone. In negotiations, the tone of a response is critical. The application frequently generated replies that were either overly defensive or excessively positive, leading to unnatural and unacceptable communication.
- Lack of generalization. The application produced effectively random responses when encountering negotiation scenarios not explicitly covered in the playbooks, limiting its utility in novel situations.
Iteration 2: feedback
Recognizing the limitations of a playbook-only approach, the team built a self-learning loop powered by direct lawyer feedback. The first feedback mechanism was simple: it saved the lawyer's decision (accept, reject, modify), the original text, the redline, and a thumbs-down signal. That provided a basic error signal but lacked important nuances.
To deepen the system's understanding, the feedback loop evolved to capture the lawyer's expected action and specific comments. When the AI gained the ability to draft modifications, the loop expanded again to collect the final modified text the lawyer actually used. This granular data proved to be a turning point: it let the system map the counterparty's intent directly to the lawyer's precise, improved language. The agent started getting much more accurate, adapting to the team's preferred style and negotiation postures.
The team also started surfacing relevant past feedback and rules directly within the suggestion interface. This gave lawyers transparency into why a suggestion was generated, building trust and letting them see exactly what influenced the AI's output.
At runtime, the agent queries a feedback store that now contains thousands of previous interactions. The system first conducts a similarity search with metadata filtering, narrowing the pool to about 20 pertinent documents. These candidates pass through a second filtering layer using state-of-the-art LLMs to verify alignment with the user's and the counterparty's intent. The agent then analyzes the most relevant examples to decide whether to accept or reject the proposed change, and generates a comment for the counterparty.
To construct the context for future prompts, the team implemented an exponential decay weighting algorithm. The mechanism prioritizes recent decisions, ensuring the agent adapts to shifting legal stances and mitigates concept drift while maintaining a balanced distribution of agree and disagree examples. This dynamic, balanced few-shot prompting lets the redlining tool learn user preferences and policy nuances in near-real-time without requiring manual model fine-tuning.
Iteration 3: tone
To address the persistent challenge of tone, the team integrated an additional LLM call at the end of the processing pipeline, specifically to modulate the output's linguistic style. Getting tone right in legal negotiations turned out to be harder than it sounds. Early outputs oscillated between two failure modes: sometimes overly defensive and adversarial, other times inappropriately accommodating and eager to please. Neither matched the measured, professional tone lawyers use in actual negotiations.
The breakthrough came from close collaboration with the in-house Legal team. Recognizing that tone is a critical component of legal communications, the engineers empowered lawyers to take direct ownership of the specific language defining the desired tone and conversational style. The engineering team managed the anatomy of the three-part prompt; the in-house Legal team managed its content.
The first part of the prompt the lawyers update is the objective. This section defines the role of the tone modulation agent, and it implements a proactive and pessimistic reflection technique by assuming that the comment generated by the agent already has the issues encountered previously. Other reflection techniques involve complex loops to converge on the desired behavior; the team found that presuming bad quality does not affect positive examples but sufficiently corrects negative examples at high reliability even with a single shot. They attribute this to reduced complexity: the LLM has one less decision to make inside a single prompt. The lack of loops also limits token usage and latency.
Next, the prompt specifies the preferred style of the lawyers. Because the agent provides the first pass of responses on behalf of the lawyer, it must be direct, first-person, and must not engage in idle conversation. Finally, a section of conversation starter sentences serves as few-shot examples written by Uber's in-house lawyers. Within days, the lawyers had tuned it to match their voice precisely.
This produced a crucial insight: prompts that directly influence output format should be crafted and managed by the application's subject matter experts or end-users, with AI teams guiding to ensure prompting best practices are met.
Iteration 4: agentic modifications
Deciding whether to accept or reject a change is valuable, but the real time-saver for a lawyer is drafting the counter-proposal. Simple generative text often hallucinates terms or drifts from the contract's defined terms. To solve this, the team moved beyond simple text generation to an agentic workflow: when a MODIFY decision is reached, the system references historical feedback and rules to refine the clause. This lets the tool help lawyers draft high-quality proposed compromise language that reflects prior legal guidance and incorporates the company's broader risk position for each particular issue.
The rules database
Finally, the team realized that lawyers frequently work with their own templates, often shared across teams. Two advantages followed. First, repeatable questions: there is significant commonality in the questions lawyers receive, and repeatedly articulating Uber's position with a consistent tone is time-consuming. Second, consistent search keys: because the original document consistently uses the same language, reliable keys can be generated for searching when a document undergoes changes.
To leverage these benefits, Uber built a rules database for lawyers. A lawyer can select important or frequently modified text within an original template document and associate a rule with it. A rule encapsulates Uber's stance, any fallback positions, and an example of the desired application response. When the application detects a change, it performs a semantic vector search of the modified sentences against the rules index. If a rule matches with high confidence, it is injected into the context window. To ensure strict adherence, a final LLM step verifies that the output aligns precisely with the retrieved specifications.
When expanding the tool to additional lines of business, the rules database provides a secondary benefit: helping lawyers maintain consistency in legal positions across the company.
Architecture
The architecture follows a thin-client pattern: the Word add-in handles document interaction and user input, while a Python back end orchestrates all AI operations. When a lawyer triggers an analysis, the back end runs a LangGraph workflow that parallelizes intent detection, risk assessment, and policy lookup. Vector stores (OpenSearch) provide the memory layer, retrieving relevant rules and historical feedback to ground each decision.
Microsoft Word add-in: meeting lawyers where they work
A critical design decision was building the assistant as a Microsoft Word add-in rather than a standalone web application. Lawyers live in Word, so asking them to copy-paste between tools would kill adoption. The add-in is built on React and TypeScript, running as a task pane alongside the document. It communicates with Word through the Office JavaScript API to read tracked changes, apply modifications, and highlight clauses under review.
The integration came with challenges. The Office JavaScript API is inherently slow, requiring optimization at the application layer with aggressive caching and minimal round-trips to Word. The tracked changes API has edge cases that cause crashes on certain document structures, requiring defensive loading strategies. And deleted text in tracked changes returns empty strings, forcing the team to maintain its own text state for accurate diff display. These constraints shaped the architecture: a thin client that orchestrates Word operations carefully, paired with a back end that handles all the heavy AI lifting.
High-performance orchestration
The team optimized for both latency and accuracy by parallelizing independent graph nodes. While the Intention Node analyzes the counterparty's meaningful intent, the Risk Level Node concurrently assesses the clause's danger profile.
Hybrid decision engine: rules versus learning
To balance flexibility with compliance, the system treats hard and soft logic differently. For non-negotiable policies, it uses a deterministic rules engine. Unlike standard RAG, this engine triggers only on strict semantic matches against a target sentence; if a rule matches, it acts as a hard guardrail, overriding the model. For negotiable nuances such as tone and strategy, the system relies on the probabilistic feedback loop. This dual approach let the team bootstrap the system, delivering high-confidence results from day one based on rules, while the feedback loop accumulated the data needed to handle more complex, unstructured scenarios.
Self-learning via feedback
The core of the self-learning capability is a feedback loop that removes the need for manual model fine-tuning. The system captures a comprehensive snapshot of every lawyer interaction. This structured dataset includes the original context (contract text, counterparty redline), the AI's analysis (intent, risk level), and, most crucially, the lawyer's full response: their final decision, any specific text modifications they applied, and their written reasoning. This granular data allows the system to model not just what was decided but why, enabling it to replicate the nuance of senior counsel in future similar scenarios.
Time-weighted self-learning
To prevent concept drift, where the model might over-index on outdated legal positions, the team implemented a custom feedback sampling algorithm based on exponential decay. When retrieving few-shot examples for the context window, the system calculates a weight for each historical interaction:
weight = exp(-lambda * age_days)
Here lambda corresponds to a configurable half-life, currently set to 365 days. This ensures the model adapts to shifting legal stances without over-indexing on outdated precedents. Effectively, the system learns the team's current preferences in near real-time, just as a human colleague would pick up on new norms.
Data storage and vector search
The knowledge base relies on OpenSearch, using nomic-embed-text-v15 embeddings for high-dimensional semantic retrieval. Two primary indices are maintained:
- Rules index: stores legal policies, vectorized by the target sentence.
- Feedback index: stores historical negotiation data, vectorized by the counterparty's intention rather than raw text, allowing the system to catch semantically identical changes phrased differently.
When a new redline is detected, the system performs semantic vector searches against these indices to retrieve the most relevant rules and past decisions, ensuring the AI has the exact context needed to suggest the change.
Improving rule and feedback retrieval accuracy
One advantage in this problem space is that negotiation always starts from the same source contract. Learnings and feedback can therefore be anchored to the same sections across multiple interactions over the same document. The system first breaks a legal contract into small chunks delimited by ;, . or newline characters. Whenever it encounters a redline, it looks up the original source sentence to retrieve guaranteed-relevant historical occurrences or explicit rules. Whenever a user provides feedback, the same verbatim source sentence is used for ingestion into the database.
Achieving symmetry between query and input embeddings matters for the best search performance. Here the stored embedding and the embedded query are guaranteed to be similar because they are all sourced from the contract. Similarity search is still required because the document may change subtly with each round of lawyer annotations.
Session memory
For the conversational interface, the system uses Redis to maintain session state, allowing lawyers to ask follow-up questions about the document. Behind the scenes, continuous LLM-as-a-judge evaluations monitor the quality of retrieved context (document_quality_metric) and the adherence of the model to company rules (rule_adherence_metric).
Today's approach: harness engineering
As part of Uber's Agentic AI efforts across the company, which include Legal AI, the team is exploring agentic harnesses such as Claude Code and OpenCode in conjunction with skills to structure and automate workflows. In parallel, it will explore incorporating legal ontologies and knowledge graphs to provide a semantic foundation for representing legal concepts and relationships, enabling more consistent interpretation and more reliable reasoning.
Conclusion
The journey through these iterations, from the initial RAG system to integrating user feedback, refining tone, and implementing a robust rules database, underscores that building a truly effective AI-powered legal assistant is an iterative process driven by continuous learning and tight collaboration with subject matter experts. Each challenge encountered became an opportunity to deepen the team's understanding and refine the system, leading to a more accurate, adaptable, and user-centric tool that genuinely augments human expertise in complex legal negotiations.
Source: Uber Engineering Blog, "Building Uber's Redlining Agent" by Austin Greco, Meghana Somasundara and Frank Tenente (uber.com/us/en/blog/building-ubers-redlining-agent, 2026-10-08). Fully translated and adapted here; commercial promotion sections from the original page were removed. All figures and the demo video are from the original article.
Source:Uber Engineering Bloghttps://www.uber.com/us/en/blog/building-ubers-redlining-agent/

