Research Agents Miss the Research - 2026-07-30
Research judgment, agent-system efficiency, accounting reliability, and real-time robot control were today's four useful signals.
In one minute
10/10 Research agents completed the engineering but missed the research
10/10 GPT-5.6 shows where agent efficiency actually comes from
9/10 Accounting agents look capable until you ask for reliability
The details
Research agents completed the engineering but missed the research - 10/10
Takeaway: Agents can execute research workflows, but still fail where judgment, backtracking, and knowing the publication bar matter.
What changed: Two six-day shadow evaluations gave frontier agents unpublished NeurIPS questions; both completed the engineering but produced work the original authors unambiguously rejected.
Sources: arXiv paper: arXiv: Can AI agents conduct open-ended AI research?
GPT-5.6 shows where agent efficiency actually comes from - 10/10
Takeaway: Agent efficiency is a systems problem: routing, kernels, speculative decoding, context control, and cache-friendly harnesses compound.
What changed: OpenAI says GPT-5.6 cut serving costs 20%, improved token generation over 15%, and tripled ARC-AGI-3 scores after harness fixes.
Sources: OpenAI: GPT-5.6 frontier intelligence and efficiency, X/Twitter @OpenAI: production efficiency results, X/Twitter @OpenAI: ARC-AGI-3 harness investigation
Accounting agents look capable until you ask for reliability - 9/10
Takeaway: A good first answer is not reliability: no accounting agent passed more than 2.6% across eight repeated attempts.
What changed: APEX-Accounting tests 160 private, expert-authored tasks across spreadsheets, PDFs, and accounting systems; Claude Fable 5 led on criteria coverage.
Sources: arXiv paper: arXiv: APEX-Accounting
TurboVLA removes the LLM from real-time robot control - 8/10
Takeaway: Real-time robot control may not need an LLM in the loop; direct vision-plus-language-to-action is smaller and faster.
What changed: TurboVLA reports 97.7% LIBERO success at 32 Hz with 0.2B parameters and 0.9 GB inference VRAM on an RTX 4090.
Sources: arXiv paper: arXiv: TurboVLA, GitHub: H-EmbodVis/TurboVLA
Quick feedback
Send what was useful, what to cut or rank lower, what was missed, and how the link list should change.



