All reports
Everything we've published
Head-to-head AI tool benchmarks and open build logs — each with the exact prompt, the evidence, and an honest verdict. Newest first.
Same invisible .env bug, three AI fixes — diff hygiene decided it
An invisible character silently breaks .env configs. All three pass the hidden gate on a real python-dotenv bug — but exit hygiene differs a lot. Full diff evidence.
Read →The perfect tool stack for a fully-automated digital-human pipeline?
6 stages, the 2026 candidates for each, our leanings — and 4 questions for people who've actually shipped.
Read →We built an AI film pipeline that wrote 437 scripts and shipped 0 films
Codex + Claude Code, stuck for 20 days. How an agent over-engineered itself into paralysis — and what we're asking for help on.
Read →Same bug, three agents, one held-out test
Codex, Claude Code and Gemini CLI fix the same real pendulum bug under an identical prompt and a hidden gate. All pass — the difference is in the details.
Read →New reports land here as we run them. Disagree with a verdict? Open an issue on GitHub.