Why a 15M-Parameter Agent Is Claimed to Beat the Craftax Final Boss 44% of the Time
A verified-scope rundown of Dan Kondratyuk's October 7, 2026 agent meta-harness claim: a 15M-parameter Craftax agent with a claimed 44% final-boss win rate, abo
Dan Kondratyuk (@hyperparticle) claimed in an X thread on October 7, 2026 that an agent meta-harness built a 15M-parameter Craftax agent in three weeks, and that the agent beats the final boss 44% of the time at a claimed cost of about one cent and 10 seconds appearing to describe a single run. Full details are on the official rekursiv.ai blog, and Kondratyuk clarified that the meta-harness itself is not open source — only its engine, trackinizer, is.

Image source: @hyperparticle / X
The core of the announcement answers a skeptical question: whether a small neural net can beat Craftax without an LLM's knowledge. Kondratyuk says experts doubted it, then presents the 44% final-boss win rate as the harness's answer. The thread evidence does not state the evaluation setup — number of attempts, seeds, or episode conditions — so 44% should be read strictly as the author's claimed figure, not an independently verified benchmark.
Meta-harness and the 4,805-experiment structure
According to Kondratyuk, the meta-harness is built on trackinizer, rekursiv-ai's open-source graph database (github.com/rekursiv-ai/trackinizer). The campaign launched 4,805 experiments, each training for a few hours on a single GPU.
The recursion he describes has three layers: agents run the research, the learner restarts from its own best moments, and each generation's games train the model that teaches the next generation. He also notes a new trackinizer UI update is free to use.
Do not confuse the two cost figures. The roughly one cent and 10 seconds appear to describe a single run of the finished 15M model, not the training bill for 4,805 experiments. The thread gives no total training cost.
Final-model play and the timeline dispute
Kondratyuk claims the final 15M-parameter model learned expert-level tactics from scratch: penning enemies into makeshift jails, pre-firing as enemies round corners, and exploiting ice-versus-fire matchups. His literal wording is that ice attacks beat fire and fire attacks beat ice — presented here as his claim, not as a verified game-mechanics fact.
The timeline wording needs care. In a follow-up post, Kondratyuk himself called the original "3 weeks" phrasing conservative: the campaign actually beat the game within a few days, and the remaining time went to ablations for a better score and to making the meta-harness discover ideas faster. Joshua V. Dillon added the same clarification: days for the result, three weeks for improving the trackinizer system.
GPT-6 Astra comparison and disclosure scope
The opening post also compares against GPT-6 Astra at a 1% win rate, taking 24 hours and $83. The benchmark conditions, execution environment, and measurement method for that comparison are not given in the thread evidence. It is therefore an author-claimed comparison, not an independent benchmark, and this article does not present it as established fact.
The disclosure scope is explicit. Asked whether the meta-harness is open source, Kondratyuk answered that they open-sourced the engine powering it — trackinizer — not the meta-harness itself. He points to the official blog (https://rekursiv.ai/blog/craftax/) for the full story. The thread also carries an open dispute: a commenter argued Craftax is an RL benchmark for general RL algorithms, not LLM-designed task-specific policies, while Kondratyuk replied the agents' improvements were fairly general, such as representing the board as a 2D grid with CNNs. No resolution appears in the acquired thread evidence.
Sources
- Dan Kondratyuk (@hyperparticle) X thread: agent meta-harness announcement
- rekursiv.ai official blog: full Craftax write-up
- trackinizer repo: github.com/rekursiv-ai/trackinizer
- Dan Kondratyuk follow-up: clarification on the "3 weeks" wording
- Dan Kondratyuk reply: answer on meta-harness OSS scope