Verification Is the Real Capability - 2026-07-20
A checkable math counterexample, a $25 WordPress exploit chain, the cost of wrong AI advice, and Kimi's capacity crunch.
A checkable math counterexample, a $25 WordPress exploit chain, the cost of wrong AI advice, and Kimi’s capacity crunch.
Essential reads
Fable produced a counterexample humans can check directly - 10/10
Takeaway: The best AI breakthroughs may be outputs humans can verify cheaply; Fable’s counterexample fits inside one tweet and two deterministic checks.
What changed: Fable helped produce a three-variable counterexample to the 1939 Jacobian conjecture; its determinant and repeated outputs are independently checkable.
Sources: X/Twitter @__alpoge__: Levent Alpoge: original Fable result, news.ycombinator.com: Hacker News discussion
A $25 agent run found and chained a WordPress RCE - 10/10
Takeaway: Give agents a falsifiable target and adversarial audit loops; one $25 run found a real WordPress exploit chain humans then verified.
What changed: Searchlight Cyber adapted OpenAI’s math prompt; Sol Ultra found a pre-auth SQL injection and chained it to RCE for roughly $25.
Sources: slcyber.io: Searchlight Cyber: complete WordPress RCE write-up, news.ycombinator.com: Hacker News discussion
Wrong AI advice erased the option to abstain - 9/10
Takeaway: AI copilots need an abstain path: fluent advice can erase uncertainty signals long before it improves accuracy.
What changed: Across five experiments, wrong AI advice reduced abstention, cut correct answers to roughly one third, and nearly doubled confidence.
Sources: osf.io: PsyArXiv: AI advice and abstention preprint, news.ycombinator.com: Hacker News discussion
Kimi’s demand spike exposed agentic inference cost - 9/10
Takeaway: Open weights do not make inference cheap: Kimi’s demand spike forced subscription pauses and separate compute pools by workload.
What changed: Moonshot paused new Kimi subscriptions after demand pushed clusters near capacity, then split general and coding plans to allocate compute.
Sources: X/Twitter @kimi_moonshot: Kimi: official capacity update, reuters.com: Reuters: Kimi subscriptions and infrastructure demand
Quick signal
The math result immediately became the benchmark joke (7/10): The joke lands because the useful benchmark is no longer a leaderboard; it is whether experts can verify the artifact.
Also worth knowing
Pretraining quality predicts returns from RL (8/10): Post-training cannot erase pretraining debt; lower loss and more pretraining tokens buy steeper returns from the same RL compute.
Xiaomi scaled robot learning with embodiment-free data (8/10): Robot learning may scale through embodiment-free trajectories first, then a smaller real-robot alignment stage.
Agent harnesses need behavior maps, not only file indexes (8/10): Agent repositories need behavior maps, not just file indexes; changes often cross modules, rare paths, and hidden state transitions.
More links
quesma.com: I burned all my tokens researching how to save tokens. (8/10 / context). A failed 111-agent run becomes a cheaper, staged multi-provider research pipeline.
OpenAI: A scorecard for the AI age. (8/10 / context). Measure useful work, successful-task cost, reliability, and scale instead of raw intelligence.
danluu.com: How I think about agentic coding and testing. (8/10 / context). Practical notes on tests, review, and process changes when code generation gets cheap.
gertlabs.com: Branched rollouts for agent evaluation. (7/10 / watch). Compare alternative continuations from the same agent state instead of grading one trajectory.
support.claude.com: Claude Fable 5 on your plan. (7/10 / context). Confirms Max-plan access and current limits for Anthropic’s research model.
Quick feedback
Send what was useful, what to cut or rank lower, what was missed, and how the link list should change.





