Agent Systems Meet Real Tests - 2026-08-06
Autonomous research, programmable harnesses, and skill compliance all moved from demos toward measurable engineering.
In one minute
10/10 HTTP Terminator turns expert security judgment into an autonomous research pipeline
9/10 Prime Agent makes context and harness state programmable
9/10 Skill-Use finds that agents trigger skills better than they follow them
The details
HTTP Terminator turns expert security judgment into an autonomous research pipeline - 10/10
Takeaway: Autonomy became useful only after PortSwigger paired hypothesis generation with deterministic validation, authorized targets, anomaly filters, and expert review.
What changed: PortSwigger released a 25-page whitepaper and AGPL pipeline that generated 30,000 HTTP-desync vectors and validated them against authorized live targets.
Sources: portswigger.net: PortSwigger: HTTP Terminator whitepaper, GitHub PortSwigger/http-terminator, X/Twitter @albinowax: researcher release post
Prime Agent makes context and harness state programmable - 9/10
Takeaway: Treat context and harness state as programmable infrastructure; Prime Agent’s gains show the wrapper can matter as much as weights.
What changed: Prime Intellect released an MIT-licensed agent with persistent IPython, programmatic subagents, durable harness state, and reviewable self-refinement.
Sources: primeintellect.ai: Prime Intellect: Prime Agent release, GitHub PrimeIntellect-ai/prime-agent, X/Twitter @PrimeIntellect: official launch
Skill-Use finds that agents trigger skills better than they follow them - 9/10
Takeaway: Skills are not instructions agents reliably obey: benchmark triggering, step compliance, and forbidden actions separately before trusting reuse.
What changed: Skill-Use evaluates 79 real skills and 177 executable tasks; the strongest model-harness pair reaches only 0.613 overall.
Sources: arXiv paper: arXiv: Skill-Use paper, GitHub JinyiHan99/Skill-Use-Bench
More links
No extra links today; the short list is intentional.
Quick feedback
Send what was useful, what to cut or rank lower, what was missed, and how the link list should change.



