Non-autoregressive RL models resurface · Microsoft director calls AI scraping the largest theft of labour · StepFun previews Step 5 · Spain blocks Archive.today
Audio Briefing
AI NEWS DAILY · 20 SEP 2026 · EN0:00 / 5:00
Headlines
01
Non-autoregressive decision models with RL top HN
A builder's retrospective on non-autoregressive decision models trained with RL, posted a year after the work. 1,238 points, 295 comments.
Microsoft director calls AI scraping the largest theft of labour
Legal briefs in the NYT lawsuit quote a Microsoft director calling AI scraping the largest theft of labour in human history, while OpenAI's head calls ChatGPT an existential threat to publishers. 171 points, 47 comments.
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with RL see one scalar reward per trajectory — sparse signal for long-horizon credit assignment. On-policy distillation from a privileged self-teacher promises dense token-level supervision, but privileged information does not guarantee a reliable teacher, and its value is stage-dependent. RetireOPD trains a skill-conditioned teacher on environment rewards first, then trains a skill-free student on RL plus distillation, with Adaptive Retirement: the student drops the teacher once their discrepancy stops shrinking. reinforcement learning, distillation, agents
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Support-agent RAG treats closed cases as static documents, which fails when the useful signal is a case's intermediate states rather than its conclusion. RAFT abstracts each historical case into a directed chain of timeline entries, retrieves at entry level, and returns the parent-case trajectory anchored at the matched state — with an optional case-level similarity graph on top. Because the retrieval layer is evaluated directly, results need no production deployment. retrieval, agents, enterprise support
MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation
LLM causal discovery rarely treats premature omission of potentially relevant relations as an explicit design objective, so candidate graphs get pruned before direct-edge judgement ever runs. MaSCoD organises candidate third variables and local structural patterns first, then judges direct edges. On Auto-MPG, DWD and Sachs — GPT-5.4 primary, GPT-4o for replication — supplying structural hypotheses before judgement beat the no-phase-one ablation on mean recall and F1 across all six dataset-backbone settings, while raising false-positive rates. causal discovery, multi-agent, evaluation
▸Terry Tao asking what human mathematicians are for and RetireOPD's Adaptive Retirement are the same question at different altitudes: when does the supervisor's signal stop adding information, and can you measure that instead of guessing. The paper's answer — track student–teacher discrepancy and retire the teacher at plateau — is a discipline most distillation pipelines hard-code as a schedule. The 1,238-point retrospective on RL-trained decision models makes the same point from practice: the interesting engineering sits in the training-signal schedule, not the architecture. Expect retirement criteria to become a first-class hyperparameter wherever dense supervision is bolted onto RL.
▸Three items attack the same bottleneck from different sides: Stagehand cutting browser-automation tokens by 80%, Cache-to-Cache passing KV state directly between models, and RAFT retrieving at the state level rather than the document level. The shared bet is that text is an inefficient protocol between systems already carrying richer internal state, and that granularity — tokens, cache entries, timeline states — is where the wins live. StepFun's Pareto-frontier preview is the counterexample: a frontier claim with no released model and no numbers is unfalsifiable, so file it under marketing until benchmarks land. The teams shipping measurable interface reductions are the ones to copy.
▸The week's legal and infrastructure news — a Microsoft director calling scraping the largest theft of labour, Spain blocking Archive.today, a weights-exfiltration site near the top of HN, RSA-896 — is one fight over what can be copied and who verifies. 'The largest theft of labour in human history' is litigation rhetoric; the exfiltration material shows the binding constraint on taking weights is operational security and logistics, not copyright. TMLR asking authors to vouch for their own papers, and MaSCoD reporting a recall gain that costs false positives, point the same way: trust comes from verifiable provenance and honest trade-offs, not confident declarations. Whoever wins the lawsuits, the verification layer — provenance, archives, cryptography — is being rebuilt in public right now.