Uber Open-Sources 'ADR': Production AI Agent and MCP Security Framework
Uber has open-sourced ADR, an enterprise security system tracking and protecting AI agents like Cursor and Claude Code. The release includes the ADR Sensor, ADR
On October 7, 2026, Uber officially open-sourced ADR (Agentic AI Detection and Response) under the Apache 2.0 license, releasing its production-proven security framework designed to track, observe, and protect employee-facing AI coding agents and Model Context Protocol (MCP) tool integrations. The accompanying research paper detailing ADR's architecture and enterprise evaluation was accepted to MLSys 2026.

Image source: Uber / GitHub
The open-source release reflects architecture validated in Uber's live engineering environment for over 10 months. Running across more than 7,200 unique endpoint hosts and handling over 10,000 daily agent sessions, the framework uncovered credential exposures across 26 distinct categories and detected 206 credentials with 97.2% precision, establishing an effective shift-left prevention layer for enterprise developers using autonomous tools.
Observability Gaps in Traditional EDR for Autonomous Agents
Conventional Endpoint Detection and Response (EDR) solutions are built around operating system telemetry, capturing low-level signals such as process launches, file modifications, and network connections. In workflows driven by autonomous AI coding assistants, however, observing downstream execution without understanding model intent leaves substantial security blind spots.
- Absence of Causal Chains: Traditional security sensors record file writes or API calls but cannot trace the underlying chain linking the user prompt, agent reasoning steps, tool invocations, and environmental context.
- Brittleness of Static Rules: Pre-configured pattern matching fails to adapt to complex agentic threats, such as multi-step prompt injection attacks or chained tool abuse across diverse business tasks.
- Inference Cost at Scale: Evaluating every single agent interaction with a high-capacity reasoning model across an entire engineering fleet incurs prohibitive latency and API expenses.
ADR bridges this gap by combining an endpoint-level telemetry sensor that reconstructs complete execution narratives with a cost-aware, multi-tiered detection architecture.
Core Architecture: The Lightweight Sensor and Two-Tier Detector
The open-source repository centers on two complementary software components: the ADR Sensor for high-fidelity data collection and the ADR Detector for scalable runtime evaluation.
- Lightweight Endpoint Telemetry (ADR Sensor): Operating with an average execution overhead of just 0.182 seconds per run, the Python-based sensor periodically inspects local caches and databases from developer tools. It parses SQLite storage from Cursor (
state.vscdb), Warp, and opencode, alongside JSONL session records from Claude Code and OpenAI Codex CLI (~/.claude/projects/,~/.codex/sessions/). By linking raw logs into coherent prompt-to-outcome sequences across macOS, Linux, and Windows, the sensor captures session configuration, active skills, permissions, and MCP tool interactions. - Two-Tier Runtime Detection (ADR Detector): To maintain high detection quality without runaway inference costs, the detector splits workload evaluation. Tier 1 executes a lightweight triage LLM optimized for high recall, filtering out obvious benign activity while flagging prompt-injection cues, credential access, or privilege changes. Flagged events escalate to Tier 2, where a reasoning agent draws on three dedicated MCP context providers—source code inspection, threat intelligence, and enterprise policy stores—to perform deep semantic threat analysis with high precision.
Two additional components described in the research paper—ADR Discovery (an endpoint inventory scanner for unapproved AI tooling) and the offline ADR Explorer engine—are not included in this initial open-source release.
The ADR-Bench Suite and Enterprise Deployment Considerations
Alongside the operational code, Uber published ADR-Bench, a comprehensive benchmark derived from production telemetry to help organizations evaluate agent defense mechanisms under realistic enterprise constraints.
- Production-Derived Benchmark: The suite comprises 303 business tasks (261 benign workflows and 42 sophisticated attack vectors), covering 17 agent attack techniques across 5 tactical stages and 133 MCP server environments (including 78 benign utilities, 25 vulnerable target tools, 12 environment emulations, and 19 community and official servers).
- Comparative Baseline Results: In evaluations against state-of-the-art baselines (ALRPHFS, GuardAgent, and LlamaFirewall), ADR achieved a 2x to 4x improvement in F1-score while detecting 67% of attacks with zero false positives. On the public AgentDojo prompt injection benchmark, it detected all attack scenarios with only three false alarms across 93 tasks.
Engineering organizations evaluating ADR should note two practical deployment factors. First, ADR-Bench is a research evaluation artifact containing synthetic attack scenarios and must be executed in isolated sandboxes (containers or virtual machines) rather than production hosts. Second, because the ADR Sensor reads local database and cache files directly from developer machines, teams must verify local file read permissions and maintain schema compatibility as upstream AI IDEs and CLIs evolve.
Sources
- GitHub: uber/ADR Repository (Apache 2.0)
- MLSys 2026 Paper (arXiv:2605.17380): ADR: An Agentic Detection System for Enterprise Agentic AI Security
- Roundtable Space (@RoundtableSpace): Uber ADR Open-Source Release Signal