Headlines
01
Releasing AI-generated mathematics, responsibly
A public release of AI-generated mathematical results, wrapped in a responsible-release policy — early norms for publishing machine-generated maths. Hours old and already 38 comments deep — the discussion outran the voting.
02
OpenAI's DevDay 2026 recap
OpenAI's own roundup of DevDay 2026 — the canonical summary of the developer-conference announcements if you skipped the keynotes.
03
Cloudflare ships Cf, an agentic CLI for its API
Cloudflare's CLI built for LLM agents rather than humans — the API surface wrapped so a model can operate it directly. Vendor tooling is starting to treat the model as the primary user.
04
First US construction permit for a BWRX-300 small modular reactor
The NRC issued the first US construction permit for a small modular reactor — GE Vernova's BWRX-300, at TVA's Clinch River site. The power half of the AI buildout just got its first permitted project.
AI / LLM / Agents
01
PSSA: a non-transformer language model from scratch in Rust
A complete language model that isn't a transformer, implemented from scratch in Rust — working code for anyone curious about architectures beyond the standard stack.
02
From bag-of-words to Jev: a history of text classification
Sebastian Raschka traces text classification from bag-of-words to Jev — a history of the NLP task that never left production.
03
Jeeves: LLM reasoning over Jev-style decision models
PostHog open-sources Jeeves, applying LLM reasoning to Jev-style decision models — the repo to read if you're wiring models into decision loops.
04
Making a vibe-coded site look designer-made
A walkthrough of the gap between vibe-coded output and designed output, and how one team closed it.
05
258 papers in nine months: the Polson number
Nicholas Polson has 258 academic papers so far in 2026 — roughly one per calendar day. The number is the story — the AI-authorship question is the discussion.
Papers — ArXiv CS.AI
01
You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
LLM routers pick a model off a static price list and stop there — the open-weight market adds a second decision the price list cannot answer: which provider serves that model. Live measurement across competing endpoints over three waves shows the same model varying sharply in quality, latency, availability and price by provider — pricier endpoints are consistently faster, but price predicts neither quality nor availability, and one deployment is near-normal on knowledge tasks yet collapses on multi-step reasoning. The paper formalises same-model provider selection as a price-taker, market-aware routing problem. [inference, routing, open-weights]
02
HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
Local hardware fits only small models, and the standard remedy for hard queries — escalating to a cloud model — gives back the privacy and cost savings that motivated running locally. HARISSA makes both remaining decisions from the model's own hidden states: whether a query is worth extra reasoning compute, and whether an answer is safe to deliver or should be deferred to a human. The prefill state, computed before a single token is generated, carries enough signal for both checks. [local-inference, safety, efficiency]
03
actr: aligning thoughts and responses for multilingual safety in reasoning llms
Reasoning LLMs can answer a non-English jailbreak unsafely even when their own chain of thought has already flagged the risk. ACTR measures that disconnect with a think gap score — how much the reasoning trace actually drives attention during generation, per language — then closes it by aligning thoughts and responses, using neuron-masking importance over a jailbreak corpus to target training. The safety reasoning already exists — the wiring from trace to response is what is missing. [safety, multilingual, reasoning]
04
EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
Enterprise and financial benchmarks test static extraction, calculation and QA, while the actual job — decisions with missing information, uncertainty, feedback and long-term trade-offs — goes unmeasured. EnterpriseBench reorganises existing QA datasets into a capability- and difficulty-annotated foundation suite and adds three interactive professional settings, starting with consulting-style client diagnosis. It scores agents across the full spectrum from static answers to dynamic, multi-step decisions. [agents, benchmarking, enterprise]
05
Diagnosing and Improving Probabilistic Reasoning in Large Language Models
An LLM acting as a decision assistant can fail twice — forming wrong beliefs from the evidence, or holding right beliefs and still choosing badly against the stated utility. This paper decomposes decision loss into exactly those two components, applies the split to frontier and open models on a synthetic benchmark with known ground truth, then runs RL interventions targeting beliefs, decisions or both across three domains to see which gains transfer and which stay put. [evaluation, reasoning, decision-making]
06
ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents
Recommendation agents that browse on a user's behalf break in two places: brittle perception over noisy, heterogeneous item pages, and inefficient long-context reasoning over extended interaction histories. ReMem replaces raw-HTML parsing with OCR over screenshots — platform-agnostic, and closer to how a human reads a page — and adds a time-evolving dynamic memory for long horizons. A concrete architecture for the perception and memory layers every user-side agent needs. [agents, memory, multimodal]
Infra / SRE / DevOps
01
NSL: WSL, pointed the other way
Show HN: NSL is WSL in reverse — a compatibility layer giving Linux native access to Windows tooling, the mirror image of Microsoft's subsystem.
02
Systems that no one will test
An essay on the code that sits outside every test strategy — unowned, load-bearing, and untouched until it breaks.
03
A staff engineer's guide to inventing work
On creating scope nobody assigned to you — the staff-engineer skill that never makes it into the job description.
Hacker News
01
NAND-16: a computer from 277,248 NAND gates
A working 16-bit computer built from nothing but NAND gates — 277,248 of them. The number is the story.
02
PS5 'Relapse' exploit published
A PlayStation 5 exploit released as a GitHub repo — 283 points, 170 comments.
03
NASA recruits ex-SR-71A staff for a secret restart
NASA has approached several former SR-71A crew and engineers about quietly reviving the Blackbird. Nobody is saying why — which is half the story.
04
Real-time solar system: 526,000 asteroids, every tracked satellite
An interactive real-time model of the solar system, tracking 526,000 asteroids plus all tracked satellites.
05
Ask HN: What are you reading?
The perennial thread at 495 comments — a reading list assembled by people who build things for a living.
Why It Matters
▸
The through-line is the model's job description changing from answer engine to operator — Cloudflare ships a CLI whose primary user is a model, PostHog wires LLM reasoning into decision models, and the papers supply the theory that framing was missing: EnterpriseBench scores long-horizon choices with missing information and feedback, while the probabilistic-reasoning work splits decision loss into believing wrongly versus choosing badly. Once a model commits rather than answers, the failure that matters shifts from a wrong fact to a wrong commitment. Most benchmarks — and most eval suites in production — are still measuring the answer engine.
▸
Delegation is outrunning verification on two fronts. actr finds models that flag a safety risk in their own reasoning and then ignore it when responding in lower-resource languages — the signal exists but is not wired to the actuator, exactly the failure you would predict once models act instead of advise. And the responsible release of AI-generated mathematics, alongside Polson's 258 papers in nine months, has publication norms being improvised in public — one 38-comment thread at a time.
▸
The economics layer is being rewritten too. The routing paper shows the same open-weight model is effectively a different product per provider — price buys latency, not quality, and one endpoint near-normal on knowledge tasks collapses on multi-step reasoning — so 'which model' is the wrong unit of procurement. HARISSA attacks from the other end: a small local model reads its own hidden states to decide when extra compute pays and when to defer to a human, keeping privacy without blanket escalation. The first US construction permit for a small modular reactor is the same buildout reaching the meter — the power layer now permitted while the software above it is still being defined.