Google Research Solves Five Open Math Problems with Gemini Multi-Agent Harness 'Cogentic'
Google Research has introduced Cogentic, a Gemini-powered multi-agent orchestration harness that autonomously solved five open problems across online learning,
On September 30, 2026, a team of seven researchers at Google Research released a paper titled "Cogentic: Multi-Agent Orchestration for Automated Proof Discovery" (arXiv:2609.40324), introducing an autonomous multi-agent harness designed to discover research-level mathematical proofs. Powered by Gemini as its foundation model, the system advanced five longstanding open research problems across online learning, auction theory, and mechanism design.

Image source: Google Research / USAnt
Moving beyond prompt engineering and single-shot sampling, Cogentic mirrors the collaborative dynamics of a theoretical research lab. By dividing proof discovery into specialized agent roles coupled with strict adversarial verification and persistent state tracking, the system demonstrates how multi-agent architectures can sustain complex long-horizon mathematical reasoning.
Beyond Single-Prompt Limits: Orchestrator, Provers, and Adversarial Verifiers
While frontier large language models (LLMs) can generate compelling mathematical conjectures in a single generation, open research problems require testing competing hypotheses, bypassing intricate technical roadblocks, and preserving intermediate discoveries across multi-day horizons. Cogentic overcomes single-shot limitations through an iterative "prove-verify" loop divided into distinct roles:
- Orchestrator: Maintains global state and dynamically allocates independent prover agents across promising proof directions and counterexample searches.
- Parallel Provers: Operate autonomously starting strictly from problem statements without requiring human hints or guiding lemmas, exploring divergent trajectories concurrently.
- Adversarial Verifiers: Evaluate candidate proofs from a skeptical posture. To prevent provers from misleading checkers with fluent but unsound arguments, verifiers operate in isolated, clean contexts. Each candidate undergoes both standalone verification and cross-draft comparison against parallel attempts in the same round to detect shared fallacies.
- Summarizers and Process Advisor: Summarizers independently brief provers on previous attempts and feedback to prevent uniform bias, while a meta-level advisor monitors cross-round logs to adjust operational parameters dynamically without injecting mathematical opinions.
Persistent Verified Ledger and State Coordination
A notorious failure mode in long-horizon AI reasoning is context drift, where models lose track of intermediate lemmas proved in prior rounds or restate them inconsistently. Cogentic resolves this through a disk-persisted Verified Ledger.
Only intermediate lemmas and ruled-out directions that successfully clear the multi-stage adversarial verification gates are written to the ledger. Once committed, downstream provers cite these verified lemmas directly as immutable building blocks. An auditor agent also reviews rejected drafts to extract and re-verify valid sub-lemmas, ensuring that partial mathematical progress is preserved rather than discarded.
Five Solved Open Problems and Practical Token Efficiency
Operating without human expert guidance during discovery, Cogentic generated complete natural-language mathematical proofs. Every result was independently reviewed line-by-line by domain experts and expanded into full companion research papers co-authored with those specialists:
- Online Inverse Linear Optimization: Established the first efficient $O(d)$ regret bound independent of the time horizon $T$, with an efficient $O(d^2)$ per-round computation cost.
- Competitive Complexity of Two-Sided Markets: Proved that adding only two sellers to the smaller side of a market allows trading revenue to match optimal allocation.
- Anytime Regret for $n$ Experts: Developed an anytime optimization algorithm whose constant factor matches that of the standard fixed-horizon formulation.
- Simple Mechanisms vs. Optimal Revenue: Improved the revenue approximation ratio for a single additive buyer from 5.2 down to 3.52.
- Price of Anarchy for Automated Bidding: Proved the optimal theoretical bound of 1.5 for two bidders and established the bound of $2 - 1/(4n+1)$ for $n$ bidders.
The framework demonstrated notable compute efficiency: most problems were solved within approximately 100 Gemini model calls, while the most complex required on the order of 1,000 calls, avoiding unbounded brute-force exploration.
Research Significance and Practical Limitations
Unlike formal automated theorem proving pipelines that rely on interactive proof assistants like Lean or Coq, Cogentic produces natural-language mathematical prose that human domain experts can immediately read, verify, and contextualize.
However, the authors note that the evaluation problems were chosen from theoretical computer science domains aligned with their own research background. Whether this multi-agent orchestration architecture generalizes seamlessly to unfamiliar mathematical fields remains an active question. Additionally, human expert review remains an essential bottleneck for final verification and publication, highlighting the continued necessity of expert-in-the-loop oversight in frontier mathematical discovery.
Sources
- arXiv: Cogentic: Multi-Agent Orchestration for Automated Proof Discovery (arXiv:2609.40324)
- Google Research Project Showcase: Cogentic Project Showcase
- USAnt Analysis on X: @USAnt_IDEA 2026-10-03 Analysis Post