OpenEnv Released: Hugging Face Turns Claude Code and Codex into RL Environments
Hugging Face released OpenEnv, an open-source capture proxy turning coding harnesses like Claude Code and Codex into RL environments without modifying harness o
On October 5, 2026, Hugging Face announced OpenEnv, an open-source capture proxy framework and TRL-based training pipeline that converts ten real-world coding agent harnesses—including Claude Code, Codex, Hermes, Pi, and OpenCode—directly into reinforcement learning (RL) environments without requiring any code modifications to the harnesses.

Image source: @ClementDelangue on X / Hugging Face
Historically, training coding agents with online reinforcement learning required researchers to reimplement an agent's internal control flow and tool execution logic from scratch as a custom gym environment. As a consequence, most models were trained inside synthetic, research-only scaffolds that nobody actually deploys in production, leading to severe performance drops when integrated into standard developer CLIs or editors. The Hugging Face research team (led by Adithya S.K., Ben Burtenshaw, and CEO Clement Delangue) bypassed this limitation entirely by introducing a transparent capture proxy that intercepts and records agent-model communication at the endpoint level.
The Discrepancy Between Benchmark Scaffolds and Real-World Harnesses
The structural configuration of an agent harness exerts an outsized influence on benchmark performance, often eclipsing raw model weights.
Evaluating Liquid AI's open-weights LFM2.5-2.6B model on identical software engineering benchmarks, the Hugging Face team revealed a stark disparity: under the same model weights, the solve rate reached 62% inside the lightweight Mini-SWE-Agent scaffold, yet dropped steeply to 33% when evaluated inside the production-grade Claude Code harness.
- Scaffold Bias: Minor variations in system prompts, tool-invocation semantics, repository context retrieval, and error return structures distort leaderboard rankings.
- Deployment Mismatch: Models fine-tuned on artificial scaffolds struggle to maintain conversational context, follow strict CLI tool schemas, or recover from unexpected system warnings in live developer terminals.
The team emphasized that training inside the native harnesses developers actually use is critical for developing models that generalize to real-world software engineering tasks.
Proxy Interception Architecture: Bridging Four Agent Protocols to vLLM and TRL
Rather than rewriting complex agent harnesses into specialized RL environments, OpenEnv introduces a lightweight, non-intrusive capture proxy.
The agent harness operates normally under the assumption that it is issuing requests to a standard LLM provider endpoint, while the OpenEnv host-side proxy intercepts and coordinates the interaction:
- Native Support for Four Major API Protocols: Bridges the primary API specifications used across coding agent ecosystems—OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Google Gemini.
- vLLM Dispatch and Token-Level Logging: Routes requests directly to high-throughput vLLM inference instances, capturing exact sampled token IDs and log probabilities (logprobs) without fidelity loss.
- Direct TRL Integration: Transmits recorded multi-turn trajectory rollouts and reward signals directly to Hugging Face's TRL (Transformer Reinforcement Learning) library to drive asynchronous policy optimization (Async GRPO).
This architecture enables ten distinct coding agent harnesses—including Claude Code, Codex, Hermes, Pi, and OpenCode—to serve as functional RL environments out of the box. Sandboxed task execution and verification are handled through isolated Harbor and Docker container environments.
Multi-Harness RL Results and 31% Tool-Call Reduction via Reward Shaping
Empirical evaluations highlighted a decisive advantage for multi-harness training over single-harness optimization.
When the LFM2.5-2.6B model was trained exclusively inside the OpenCode harness, its solve rate within OpenCode surged from 34% to 58%, yet gains failed to transfer meaningfully to other agent environments. In contrast, when trained simultaneously across four distinct harnesses, the average solve rate climbed from 42% to 54%, with Claude Code performance improving from 33% to 49%.
- Advantages Over Supervised Fine-Tuning (SFT): An SFT baseline trained on 3,189 high-quality rollouts generated by Qwen3.8-27B plateaued at a 47.5% solve rate, underperforming both RL regimes. Static offline trajectories do not teach compact models how to react dynamically when a harness returns execution errors following an imperfect tool call.
- Reward Shaping for Tool Efficiency: By introducing a compact reward bonus for completing tasks with fewer tool invocations, the model learned more direct problem-solving trajectories. Across already-solved tasks, tool calls dropped by an average of 31% across all harnesses, and by approximately 50% under the Codex harness, substantially curbing execution latency and compute costs.
Open-Source Release Assets and Infrastructure Considerations
To facilitate community experimentation and reproducible research, Hugging Face has released the entire project stack under open-source licenses.
The release comprises the OpenEnv capture proxy codebase, TRL trainer integrations, benchmark task datasets, the 3,189-rollout SFT corpus, complete training configurations, and seven pre-trained checkpoint weights representing progressive training milestones.
Practitioners looking to deploy local rollout loops and async GRPO pipelines will need sufficient GPU compute resources to concurrently host vLLM inference and orchestrate containerized sandboxes (Harbor/Docker). The research team also cautioned that training against a single harness risks overfitting to narrow tool-call dialects, underscoring the necessity of multi-harness curricula. Hugging Face plans to follow this initial release with larger-scale training runs applied to higher-capacity open foundation models.
Sources
- GitHub: huggingface/OpenEnv
- Clement Delangue Official X (@ClementDelangue): OpenEnv Multi-Harness RL Release
- Adithya S.K. Research Thread (@adithya_s_k): The Ultimate Guide to Multi-Harness RL