Forty-two stories cleared 200 points in today’s scan — the top 160 by rank plus a 40-hour sweep for anything that had already fallen off — and thirty-two of them are new. The front page is led by a model launch that is not a launch: Gemini 4 Argon is gated to “trusted cyber defenders” through Google’s Fairwind Program, with a price list published anyway.
The new crop is unusually infrastructural. The day’s most consequential story is not the model but a cache — a 437x compression of KV memory published by a Chinese lab, already visible in the cache-read prices of Opus 5.5 and GPT-6.1 Sol. Around it: a Microsoft token bug written up by a 16-year-old whose post Microsoft edited before publication, a phone-forensics vendor making claims through a leaked promo video, and the front page’s recurring argument about who gets to distribute software.
Gemini 4 Argon — 1,643 points
1,643 points · 1,127 comments · blog.google · HN discussion
Google’s newest frontier model, announced at the end of September and leading the front page today with the largest score of the month. It is rolling out first to “a set of trusted cyber defenders” through the Fairwind Program, with developers, enterprises and consumers to follow “as soon as possible.” The headline capability is a 1M output token limit, up from 64K — enough headroom that a single trajectory can generate hundreds of thousands of tokens of thinking. Introductory pricing is $2 per million input, $10 per million output, with cached input at 95 percent off input pricing, which is now the standard shape of a frontier launch.
The evidence offered is a stack of benchmark wins: 77.9 percent on DeepSWE v1.1 for long-horizon software engineering, first place on the Vals Index (economic impact across finance, coding, legal and tax work, weighted by each sector’s share of US GDP), leading on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark, first on Zapier’s AutomationBench at 51.3 percent, and 91.7 percent on LVBench for long video understanding. Google also says Argon is strong enough at offensive cybersecurity to be released to trusted defenders without cyber guardrails.
The most interesting numbers in the post are internal and unauditable: agents migrating C and C++ codebases to Rust at Google, from tens of thousands of lines in libraries like re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel, all under automated and manual audit before production. For libgav1 the agents replaced 32K lines of hand-written SIMD with safe Rust the compiler would vectorize on its own, producing a memory-safe decoder 2.7x faster than the existing Rust port. Separately, a team of Argon agents read fleet-wide profiling telemetry and freed 300 TiB of memory, with 500 TiB to 1 PiB projected; quantum researchers got a 40 percent improvement over a published baseline on a subroutine bottleneck. Set against that: anybody outside Google and Wiz can neither run the model nor reproduce the benchmark, the model that leads the GDP-weighted index is not purchasable, and the “introductory” price list is a promise about a future that has not been dated. The comment thread’s best framing is that the real announcement is the 800,000-line C++ to Rust migration, which is a claim about internal tooling that no competitor can match or verify.
Pi 1.0 — 806 points
806 points · 285 comments · earendil.com · HN discussion
Earendil marked Pi as stable, after “hundreds of thousands” of weekly users and a year of accumulating feedback. The 1.0 feature list is: native Codemode support (MCP plus non-LLM models like Jev and image models), extension support for virtual models, deferred tool loading, cache warming for Anthropic models, mid-conversation system messages, a new TUI theme, and full-screen by default. The pitch is that everything on that list survived being thrown at a wall for months, and the list of things that didn’t stick is longer.
The post is careful about the difference between adopting a protocol and deciding it is good — the same team published “You said no MCP” two days ago and explained that they supported MCP because the interpreter they were building anyway made it nearly free. Here the same logic recurs in the comments and lands better than the announcement does: a user with a weak laptop notes that Pi was the only harness that ran local models worth a damn because it does not ship a gargantuan system prompt, and another asks why cache warming for one vendor’s models has to be bundled into the minimal agent rather than living as a package. That question is the one to watch. Every “minimal” harness that has shipped this year has grown provider-specific plumbing, and each addition is individually defensible.
StreetComplete on iOS is now in public beta — 529 points
529 points · 134 comments · github.com/streetcomplete · HN discussion
StreetComplete is the app that turns OpenStreetMap data maintenance into simple questions — is this street lit, does this crossing have tactile paving, what are the opening hours — answered by people walking past. It has been Android-only for its entire life, and the master ticket coordinating an iOS port was opened in December 2023. It is now in public beta on TestFlight.
The engineering answer to “why did this take three years” is in the ticket: the codebase is 100 percent Kotlin, so the port uses Kotlin Multiplatform for logic and Compose Multiplatform for UI, which is not a port at all — the entire interface has to be re-created incrementally in Compose. The funding answer is more interesting than the technical one: the German Federal Ministry of Education and Research paid for part of it through the Prototype Fund, and NLnet paid for the rest. Public map data maintenance, funded by public research money, on a platform whose owner charges for distribution. The comment thread supplies the counterweight you would expect and one you would not — a user describes making real edits and then getting reverted by pedants arguing over tagging conventions, which is the known social failure mode of crowd-sourced map data, and the practical detail that the beta invite link was buried three clicks deep in the issue.
Singapore govt dating app uses Gale-Shapley stable marriage algorithm — 445 points
445 points · 462 comments · twitter.com/tuakdotsol · HN discussion
FirstDate, Singapore’s state-run dating app, runs Gale-Shapley — the 1962 stable marriage algorithm that won the 2012 economics Nobel and still assigns medical residents and kidney donations. The implementation: you build a ranked list from preferences and dealbreakers, proposers offer down their list until nobody wants to swap, and the result is stable in the technical sense — no two people both prefer each other to their assigned match. One match per cycle, a 72-hour decision window, contact details revealed only on mutual yes, Singpass verification so everyone is who they claim. Public servants aged 21 to 35 for now.
The interesting part is not the algorithm, which is sixty-four years old and well understood, but the incentive and the measurement. A commercial app is paid to keep you swiping and can only observe whether you stopped opening it; a state with marriage records can measure the outcome it says it wants. The comment thread gets to the real weakness quickly and from several directions: people do not reliably know their own preferences, the preference categories that dating apps collect mostly do not predict compatibility, and Gale-Shapley gives one side the best possible stable outcome and the other side the worst among those it would tolerate — so who proposes is a policy decision, not a detail. One commenter asks which version Singapore implemented and gets no answer.
Clef: Open-weight decision models, and new RL fine-tuning platform — 443 points
443 points · 163 comments · blog.cloudflare.com · HN discussion
Cloudflare released two decision models, Clef and Clef-flash, on Workers AI. A decision model is the narrow, cheap, deterministic counterpart to an LLM: you hand it a support message and it returns typed answers with probabilities — urgent or not, which team — so your code can route, escalate, or defer to a human without a model reasoning freely in the loop. Clef is Jev-API compatible, leads the Jev Decision Index, and the weights are on Hugging Face under Apache 2.0. There is also a new reinforcement-learning product for fine-tuning Clef on your own labels.
Two things in the comment thread matter more than the announcement. First, the licence: open weights on top of proprietary Qwen starting points, with the data and training pipeline unpublished, is not open source, and the top-voted correction is that weights are not source. Second, the price arithmetic — Jev at $0.042 per million input tokens versus Clef at $0.24 per million input with no output price listed, which at roughly 300 tokens a call puts a million decisions at about $12.60 against $72. That comparison assumes the two are interchangeable, which is the claim the benchmark is supposed to establish, and the benchmark belongs to the company whose model Clef beat within weeks of the paradigm being published. Nobody should be surprised that a leaderboard two weeks old has a new leader; the question is whether the task — bounded classification with typed output — is one where a lead is stable.
The AI Race Just Got Awkward — 409 points
409 points · 455 comments · insufferable.dev · HN discussion
A short essay arguing that the distillation narrative is dead and has been replaced by something the Western labs do not want to name: adoption. DeepSeek published MLA, which compressed the KV cache roughly 15x, then Compressed Sparse Attention, then Heavily Compressed Attention, and DeepSeek-V4.1-Flash now closes at 890 bytes of KV cache per token by combining cross-layer cache reuse and FP4 caching. Against DeepSeek-V1’s 389.12 GB per million tokens, that is 437x less memory — 99.77 percent. Serving long contexts is a VRAM problem, so this is the cost structure of every coding agent in the world.
The essay’s evidence that adoption actually happened is prices rather than rhetoric. Anthropic’s cache read on Opus 5.5 is 60 percent cheaper than Opus 5’s, and OpenAI’s cached input on GPT-6.1 Sol is 80 percent cheaper than GPT-5.6 Sol’s late-July pricing, in both cases with cache write following. Both launches were quiet, with little pre-announcement, and the author reads the silence as embarrassment. Where the piece is weakest is the question at its centre: it asks repeatedly why Chinese labs would give this away and offers no answer, then treats the fact as a gift. The comment thread supplies the answer the post avoids — commoditize your complements, especially when your comparative advantage is manufacturing — along with the more cynical read that Anthropic’s distillation complaints are about regulatory groundwork rather than theft, since Anthropic’s own models were trained on other people’s work. One commenter reports running Qwen 3.8 27B locally on a 5090 at 170 tokens per second; that model was frontier-tier nine months ago.
Why the Bronze Age Collapsed — 393 points
393 points · 290 comments · worksinprogress.news · HN discussion
Patrick Fitzsimmons, in Works in Progress, asks the better version of a well-worn question. Not why the Bronze Age empires fell — drought, famine, invaders, the Sea Peoples — but why nothing like them came back for centuries. His answer is material and political at once: bronze required tin, tin came from as far away as Afghanistan or Western Europe, and that supply chain kept weapons concentrated in the hands of palaces and elites. Iron could be made locally. Once ordinary Hittite and Egyptian farmers had iron, imperial centres no longer held a structural advantage over the people they ruled, and the political equilibrium had changed for good.
It is a good argument with an obvious weak seam: catastrophes in the ancient world were frequent, and a single technological variable doing all the work of explaining centuries of non-recovery is the same monocausal move the essay opens by criticising. The comments pull on the thread from opposite ends. One correction is that ordinary people were not using copper tools either — pure copper is too soft for a blade — which puts a dent in the essay’s neat bronze-versus-iron stratification. Another is that the Spanish did not beat the Aztecs with iron: they were vastly outnumbered and won because they had indigenous allies. The strongest comment is not a correction but an extension, pointing at Eric Cline’s palace economies, where royals literally owned the land and the people on it, and asking how much of the collapse was bad seasons and how much was a system with no slack. The modern parallel the essay draws is asserted rather than demonstrated, and the comment thread notices that the tools of violence are concentrating again in exactly the way the piece ends by warning about.
A brief history of the Bloomberg terminal — 377 points
377 points · 168 comments · spectrum.ieee.org · HN discussion
IEEE Spectrum’s October issue runs a history of the machine that trained a generation of finance into a keyboard. The piece is member-gated past the opening, so this entry is assembled from the available text and the thread. The origin facts are that IMS — Bloomberg’s predecessor — had exactly one customer, Merrill Lynch, which put in $30 million (about $110 million today) for 30 percent of the company and five years of exclusivity, waived in 1984. The first 22 Market Master terminals went to Merrill in 1982, in the middle of a global recession, and terminals have been leased in two-year cycles ever since.
The comment thread knows the machine better than the article can. The modern Terminal is a private fork of Chromium dressed to look and behave like a VT100, packaging Bloomberg’s own networking and security; it predates HTTP and backwards compatibility is treated as the product, to the point that the company keeps a museum of second-generation hardware so old workflows still have somewhere to point. A former avionics person makes the design point that matters: the density and terseness of a terminal display — everything needed, nothing more — is the same discipline as a cockpit, and both are hostile to redesign. The missing piece the thread flags is that Reuters had a competing terminal, and links two archives that treat the same history from the other company’s side.
Returning from vacation? The government can search your phone without a warrant — 356 points
356 points · 327 comments · arstechnica.com · HN discussion
An immigration advocate is suing border agents who demanded his phone, which is a small case about a large exception: the border-search doctrine means warrants are not required at the boundary, and the practical consequence is that re-entering your own country is the moment your devices have the least protection. The Ars page is mostly a privacy-preferences wall, so the substance here lives in the thread.
The thread’s two worthwhile observations are not about the Fourth Amendment. One is jurisdictional arbitrage: agents on one side of a border have powers they cannot exercise a few miles inland, so investigations are timed to the crossing — a pattern, one commenter says from having worked the intelligence side, that is not limited to the United States. The other is that the more troubling defect is procedural rather than constitutional, because a person can be searched and never told what the search was for. Quoting the amendment back at each other is the ritual here; the accountability gap is the part that does not get litigated.
Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026 — 333 points
333 points · 384 comments · techpowerup.com · HN discussion
On Micron’s fiscal 2026 earnings call, CEO Sanjay Mehrotra told investors that memory and storage supply will be tighter in fiscal 2027 and 2028 than in an already-tight 2026, that more than 75 percent of fiscal 2027 shipments are already committed, and that customers will pay much higher prices than they paid this year. Twenty-six strategic take-or-pay agreements are in place, which is a way of describing both a backlog and a floor. The TechPowerUp page is behind a bot check, so the numbers here come from the earnings call coverage at CIO, The Register and The Stack.
An earnings call is a price signal as much as a supply forecast, and the comment thread treats it accordingly: a shovel seller announcing that shovels will be expensive in winter. The more serious critique is strategic rather than moral. If you own most of the market and choose to under-supply into demand that is being created by national-security-adjacent spending, you are also funding the competitive capacity that gets built to fill the gap — and take-or-pay contracts make the shortage contractual rather than incidental. The counter-argument in the thread is worth keeping: AI demand is genuinely hard to build capacity for, and capacity built for a peak nobody has yet sized is how memory makers went bankrupt last time.
The last time my family was replaced by technology — 322 points
322 points · 624 comments · manuel.darcemont.fr · HN discussion
A French developer writes about his great-great-grandfather, a farrier in Mandres-en-Barrois who shod horses and repaired farmers’ carts, saw a car, and became a mechanic — as every man in the family after him did. The argument is that developers waking up worried about their jobs are the same person one century later, and that the farrier traded his horses for engine grease and kept the purpose: helping people get around.
The author shows up in the comments to head off the obvious misuse of his own post, saying explicitly that it is not a “shut up and adapt” lesson and that he was not trying to dismiss anyone’s anxiety, which is more honest than most versions of this essay. The thread’s best pushback is structural rather than emotional. A farrier in one village served local demand that a machine could not serve from elsewhere, and the machine that replaced him did not improve itself weekly or replicate itself at zero marginal cost. The comparison that lands hardest is against horses: there is no rule that better technology creates more or better jobs for horses, and it sounds absurd to say out loud — substitute humans and it sounds roughly right. The interesting tension in the thread is that everyone agrees the farrier analogy is comforting in the wrong direction and nobody has a better one.
I could’ve accessed 17T Microsoft records — 313 points
313 points · 127 comments · blog.faav.net · HN discussion
Faav, a 16-year-old security researcher, found that an internal Microsoft analytics service never checked the signature on a login token. That let him claim an administrator’s identity and submit unauthorized SQL queries, reaching an estimated 17.3 trillion stored rows across a wide range of datasets. He used table descriptions, metadata and bounded sample rows to size the exposure, reported it, and did not touch customer data. The bounty was $5,000.
Two sentences in the write-up matter more than the vulnerability. The first is that the impact is hypothetical — what an attacker could have done — which is the correct way to frame a finding you did not exploit, and also the reason bug bounty economics stay depressed. The second is that Microsoft had editorial control over the post and used it: sections and figures were cut, and the description of impact was reshaped before publication. The thread is nearly unanimous that the second admission is the more damaging one, because it establishes that the vendor gets to edit the record of a security failure it wishes to downplay. The counter-argument, that coordinated disclosure agreements routinely give vendors review rights, is also the reason nobody should treat a vendor’s own vulnerability write-up as neutral.
Fuck Android Developer Verification Program — 311 points
311 points · 131 comments · twitter.com/0xcrypto · HN discussion
A developer’s account of a run of administrative dead ends, which by the end of 2026 is the normal route for anyone shipping small Android software. His Play Store account was closed for inactivity while he was not working on the game he had started. He could not restore it and found no way to create a replacement. He decided to skip the store and ship an APK through F-Droid instead — and found that distributing outside the store now also requires registering for the Android Developer Verification Program: $25, plus identity verification, plus a declaration about his publishing intentions. Even development builds passed to a handful of friends, and even developers under 20, have to be registered and payment-profile verified.
The thread generalises the complaint fast and with receipts. Multiple commenters had apps published and lost their accounts for inactivity, with no restoration path and no way to stop the emails because they can no longer log in. One got a Google survey asking which Play issues frustrate developers and screenshotted an option written by Google itself: sudden account termination without actionable feedback or clear restoration guidance. Another describes the sole-trader path in full: to publish at all you need twelve testers who join a Google Group and opt in to a beta, which is a distribution requirement disguised as an anti-abuse measure. The substance of the criticism is not the $25. It is that sideloading — the escape hatch that made Android’s openness real — now routes through the same identity and payment verification as the store, and that verification is one-directional: developers verify, Google does not explain.
LinkedIn Larpmaxxing — 304 points
304 points · 240 comments · hereticpleb.vercel.app · HN discussion
An attack on LinkedIn’s feed culture that is funnier than it needs to be. The author catalogues the genre — “my third cousin got hit by a truck today, here’s what it taught me about business management” — then inspects the projects that dominate the feed and finds the same hand-tracking computer vision demo, over and over, almost certainly vibecoded from the same tutorial. The follow-up project is a detector for AI slop in posts, which means the feed now has an arms race between generated posts and generated detection.
The comment thread turns out to be about the platform’s actual function rather than its culture war. One user built a job-search agent and found the site’s JavaScript deliberately obfuscated against automated browsing, which is a strange posture for a site whose product is public professional information. The defence is the one that survives every round of this argument: it is a digital rolodex, and the people who got jobs through it got them from a searchable archive of fifteen-year-old conversations, not from the feed. The post’s own joke is that everything on it is performance, including the post; the thread’s version is that the useful layer is a contacts database with a hallucination engine bolted on top.
The top secret URSALA, RAQUEL, and FARRAH satellites (2025) — 298 points
298 points · 158 comments · thespacereview.com · HN discussion
Dwayne Day’s Space Review piece on the hitchhiker electronic-intelligence satellites, which started in 1963 when the Air Force launched the first small payload off the side of a larger satellite and ran for more than forty years under names that were classified until recently: URSALA, RAQUEL, FARRAH, GLORIA, CARRIE. The satellites were about the size of a large suitcase, covered in antennas, spun rapidly in orbit, and swept the ground for radar and other emissions, recording signals for later transmission to a ground station. Program 989 flew them as subsatellites of the HEXAGON photo-reconnaissance satellites — twenty launches from California between 1971 and 1986, one failure. It is a 2025 article resurfacing via a related submission about a satellite named after Farrah Fawcett breaking up in orbit.
The comments are better than the article in one direction and worse in another. Better: the pointed observation that the National Reconnaissance Office had put a large pile of declassified internal histories online unsorted and barely legible, which is where several of the documents in the piece came from, and that in 2012 the NRO handed NASA a pair of surplus telescope assemblies that turned out to be Hubble-class optics pointed the other way. Worse: the thread is thin on what any of this says about current capability, which is fair, because the material declassified is fifty years old and the names being revealed are the point of the piece rather than a revelation about the programs that replaced them.
RIP, vector database — 286 points
286 points · 78 comments · turbopuffer.com · HN discussion
turbopuffer announced a storage rearchitecture and used it to declare the category it helped define over. Version 1 was a serverless vector database: object storage as the source of truth for cost, tiered NVMe and memory caches for speed, customers like Cursor and Notion validating the trade. Version 2 added text and regex search strong enough that Linear uses it as a syncing engine, and mostly kept the storage layout — an ANN index as the primary index, with everything else keyed to it. That is the part going away: write amplification from maintaining an index that owns document placement has hit diminishing returns on indexing throughput, and v3 changes how documents and indexes are laid out, written, compacted and queried, with the aim of making text, regex and vector search all faster and moving more SQL to the same engine.
The founder’s comment in the thread is the honest version of the marketing: vector databases were always more about retrieval than about vectors or storage, and the term stuck too long. A competing implementation pushes back mildly by existing — LanceDB already treats the ANN index as secondary, with rows in fragments that the vector index never moves — which suggests the architectural insight is not proprietary so much as necessary, and that the differentiator will be operational. What to be sceptical about is timing: v3 is announced as coming, benchmarks are not published, and storage rewrites are the kind of change that reads well and hurts in production for six months.
Before pixels: Modular industrial dashboards — 284 points
284 points · 52 comments · unsung.aresluna.org · HN discussion
Marcin Wichary photographed modular industrial control panels in German and Polish museums and writes about them as design systems rather than nostalgia. There is an air traffic control display at the Deutsches Museum in Munich and a Fernmeldemuseum Stuttgart array of the modular panel units, and the observation that these are interaction systems, not just displays: fixed metal, fixed labels, and a physical grammar that operators learned once and never renegotiated.
The thread explains the operation of things the photographs can only show. The Spurplan track-plan panel sets a route by holding the button at the origin and pressing the button at the destination, with conflicts resolved by the panel rather than by the operator — a two-button protocol for a constraint solver built out of relays. A commenter who watched a 1990s documentary inside FedEx’s operations centre describes velcro flight strips on a wall, where moving a plane meant physically moving a card to a new time slot. That is the design argument the post does not quite make explicit: physical dashboards encode state with affordances you cannot accidentally reconfigure, while software dashboards can be rearranged and relabelled, which is why the same information is harder to read on a screen that can display anything.
56k.rip – the 1996 dial-up internet experience — 260 points
260 points · 108 comments · 56k.rip · HN discussion
A browser simulation of dial-up that takes the constraint seriously rather than cosmetically. You pick 14.4, 28.8 or 56k, and each page is delivered against a byte budget at that rate, so the same page really does take longer on the slower line; photographs resolve line by line, downloads crawl and mostly fail, and sooner or later somebody picks up the extension and the call drops. The handshake sound is a recording of a real modem, credited, while most other sounds are synthesized from a written score. Behind the modem is a small 1996 service to read: a fishing page, a recipe page, a guest book, a webring, classifieds, missed connections, a weather report, a television grid, and an article explaining what the internet is, plus mail, a chat lobby, minesweeper, cards, paint, a command prompt and screen savers.
The complaints in the thread are the correct kind of complaints for a project like this. The icons do not look like the era, the reconnect sound is wrong, and pages load too fast — which is to say the fidelity is in the simulation and not in the period detail, and the people who remember 1996 can tell. The interesting engineering question, asked and only partially answered, is where the compute goes for minesweeper in a browser pretending to be a 486, and a related project shows up in the replies with an entire virtual network stack and a Basic-alike, which turns the toy into a testbed.
Book of Shapes – Collection of minimal, generative and customizable SVG-patterns — 257 points
257 points · 23 comments · bookofshapes.com · HN discussion
Nikolaj Sokolowski’s catalogue of generative SVG patterns, browsable by family — grid, radial, noise, flow, isometric, organic, distortion, physics — with per-pattern customisation and popular/liked sorting. Everything is procedurally drawn and exported as vectors, which makes it useful as a starting point rather than a finished asset library.
Twenty-three comments make it the quietest story on the page today, in the same slot Phyllotaxis occupied yesterday: a visual tool with an obvious craft floor and nothing to argue about. It is worth noting as a pattern rather than a story. Two days running, the least contested item on the front page was a page of generated shapes, which after a day of benchmark claims and licence disputes reads like the only artefact nobody wanted to litigate.
OpenDLSS: A Vulkan Reimplementation of Nvidia’s DLSS 5 Neural Rendering Network — 249 points
249 points · 117 comments · github.com/maanHimself · HN discussion
Someone rebuilt NVIDIA’s DLSS 5 neural rendering network in Vulkan, bit-exact against build 310.8.0 — not the final image, but all 75 block boundaries, byte for byte. The architecture is a U-net of shifted-window transformer blocks with a global ViT at the bottom: 71 blocks across six pooling levels, FP8 E4M3 activations with FP16 accumulation, 141 MiB of weights. There is a second independent implementation that runs in a browser via WebGPU with no tensor cores and no FP8 at all. It takes a rendered frame — an LDR proxy, three lanes of Gaussian noise, the previous frame’s output reprojected, five conditioning scalars — and produces an RGB residual plus a temporal-blend logit. It is a generative re-renderer, not an upscaler, and NVIDIA described the model in its own report.
The legal substance is in one line of the README: you supply the weights. The code is open, the network graph is documented, and the weights — the part NVIDIA actually licenses — are not distributed, which is why the project can exist at all. Bit-exactness against a fixture you also have to supply is a strong claim that is hard for a third party to check, and the repository is four commits old with under a thousand stars, so treat the parity claim as the author’s until someone reproduces it. What is not in doubt is the direction of travel: the last two generations of hardware-locked features have been reverse-engineered within months, and the moat keeps turning out to be the weights, the driver, and the store.
Meta Uses A.I. Data Centers to Avoid Billions in Federal Taxes — 249 points
249 points · 235 comments · nytimes.com · HN discussion
The Times reports that Meta cut its 2025 federal tax bill from $9.6 billion to $2.8 billion — a 71 percent reduction — largely by classifying AI data centers and the chips in them as experimental, which makes the hardware eligible for the research and experimentation credit’s supplies rebate. The credit was created in the 1980s to spur innovation and applies to supplies used in experimental efforts rather than standard business operations. Meta’s claimed benefit from the credit went from roughly $700 million in 2023 to about $3.9 billion in 2025, which would make it the largest publicly traded beneficiary of a credit Congress wrote for somebody else. Senator Warren sent Zuckerberg a letter on the subject on 27 September, and ITEP published its own critique on 30 September. The Times page returns 403 to automated requests, so the numbers above come from the Yahoo Finance summary, the AI Weekly alert, ITEP and the Senate letter itself.
The comment thread is unusually good on the accounting. The defence is narrow but real: training a model is genuinely experimental, and AI data centers are dual-use, so some fraction of the capex is arguably in scope. The problem is that the classification is applied to the whole facility, including the inference fleet that runs as an ordinary business, and that the credit’s own test is whether the supplies were used in experimentation rather than in operations. The best line in the thread is an experiment nobody has run: ask the Times’ own accountants how they classify their engineers for the same credit. The broader point is that a credit designed for research has been captured by capex, and whether that is legal is a question nobody has litigated because nobody has been charged.
Surprisingly complex waves reveal the brain’s inner workings — 248 points
248 points · 95 comments · quantamagazine.org · HN discussion
Quanta on traveling waves, which were long treated as the hum of a working brain — the collective electrical activity of neurons, measured by electrodes, propagating like a stadium wave — and are now looking like something closer to a control mechanism. An April 2026 paper in Nature Communications, using intracranial electrodes in humans, reports a menagerie of patterns: source waves emanating from a point, sink waves converging on one, and vortex-like spiral waves, with different behavioural tasks associated with different patterns. The claim, quoted from MIT’s Earl K. Miller, is that the field has moved from “are they relevant” to “this is a major motif of how the cortex processes information.” The reason it matters: neurons take days or months to change their connections, while behaviour has to adapt in seconds, so something larger and faster has to be reorganising the network in real time.
What is solid here is the observation of the patterns, in humans, with good spatial resolution, which is a genuinely hard experiment. What is soft is the function. “May help reorganise” and “a significant number of neuroscientists now believe” is the language of a field that has found a correlation and is reaching for a mechanism, and the sample geometry of intracranial work — electrodes in patients who already need surgery — constrains which tasks and which regions can ever be tested. The football-crowd analogy does more work in the piece than the data does, which is normal for Quanta and worth flagging anyway.
Cops Can Bypass iPhone’s Automatic Reboot to Get into Locked Phones — 246 points
246 points · 201 comments · 404media.co · HN discussion
404 Media obtained a promotional video in which Magnet Forensics, which makes the GrayKey unlocking device, claims it can freeze iPhones in a state that makes them easier to break into — specifically, defeating the inactivity reboot Apple added to iOS after 404 Media’s own November 2024 reporting, a feature designed to push a seized phone from “locked once” back to “never unlocked” so that forensic tools face the stronger encryption state. That is the whole story: a vendor’s capability claim, delivered as leaked marketing, reported by the outlet the feature was built to frustrate.
Everything about it should be read as an arms-race dispatch rather than a vulnerability disclosure. There is no CVE, no technical detail, no independent reproduction, and the artefact is a sales video for a product sold to police forces, including in countries whose interest in unlocked phones is not judicial. That does not make the claim false — Apple and the forensic vendors have been trading moves for a decade, and each side’s announcements are marketing — but it does make it unauditable, which is the same problem as the Fairwind-gated model at the top of this page and the vendor-edited write-up a few sections above.
Pi Durable — 245 points
245 points · 27 comments · earendil.com · HN discussion
Shipped the same day as Pi 1.0 and explicitly not a replacement for it: an experimental package for agents that run for a long time, anywhere. The requirements list is the interesting part, because it is the opposite of the coding-agent brief. Runs on any surface, supports conversations that never end, survives catastrophic internal and external failures, and lets multiple humans steer the same agent — as against Pi’s model of one person driving one terminal and restarting by hand when the process dies.
One thousand and fifty-one points across two posts, one saying the software is now stable and built to last and the other saying the frontier is durable, reachable, multi-user agents. Both can be true, and the second is the harder problem: durability and multi-human steering are distributed-systems concerns, and the harness that solves them will look much less like a terminal and much more like a job queue with a chat interface. The 27 comments are nearly all people who use Pi and appear to be sceptical for the right reason — the last two harness features announced this week (cache warming, virtual model extensions) are plumbing, and plumbing is what makes minimal tools stop being minimal.
What TLA+ can and can’t check — 240 points
240 points · 47 comments · buttondown.com · HN discussion
Hillel Wayne, a TLA+ educator and evangelist, asking everyone to calm down after Boris Cherny mentioned that Opus found race conditions with TLA+. His objection is not to the result but to the narrative that formal methods will solve agentic software development, and he makes a sharper point than the usual “designs aren’t code” complaint: to verify a property, you must first be able to express it. TLA+ gives you invariants (always P), next-state and action properties (P’), and eventualities (P eventually), and the class of bugs it cannot find is not limited by compute or by the tool’s semantics — it is limited by properties nobody thought to ask about.
That leaves the interesting question open in a way that flatters nobody. A model that can write a specification is useful precisely because writing the specification is the part humans skip, and finding a race condition in one is evidence that the specification was adequate to the bug, not evidence that specifications are now solved. The 47 comments are the standard mix of “formal methods people have been saying this for years” and one genuinely good observation about the pipeline problem: the model can also write a specification that is wrong in a way that makes the implementation look correct.
EDG C++ front-end goes public — 240 points
240 points · 125 comments · edgcpp.org · HN discussion
On 30 September, the source for Edison Design Group’s C++ front end went public, with The C++ Alliance as its nonprofit home. EDG has been the only production-quality source-to-source C++ engine of its kind for thirty years, and a large fraction of the industry’s C++ tooling — compilers, static analysers, IDEs, and anything that needed to parse C++ properly before it needed to emit code — has been built on it under licence. John Spicer, who wrote much of it, published a note on what comes next.
The framing on the site is continuity: a change of steward, not a change of course, with fiscal sponsorship, professional maintenance and open contributions. That is exactly the claim to be sceptical about, because the failure mode of an open-sourced commercial component is a nonprofit that inherits the maintenance burden without the revenue that funded it, and the site’s own history section is still a placeholder saying the story is being written. The counter-argument is that thirty years of accumulated conformance work is not something a community can replace, so the alternative to a funded nonprofit is worse: a proprietary front end with one ageing author and no successor. This is the kind of infrastructure event that gets 240 points and matters ten years from now.
How to speed up the Rust compiler in September 2026 — 233 points
233 points · 118 comments · nnethercote.github.io · HN discussion
Nicholas Nethercote’s two-monthly audit of Rust compiler performance, and the numbers are good: a mean wall-time reduction of 4.57 percent between 29 July and 28 September, with 555 of 629 benchmark measurements improving and only 74 regressing. He calls it a sea of green. The individual wins include rustdoc improvements, profile-guided optimisation for Clippy from Jakub Beránek with an 18 percent best case, and an LLVM 23 upgrade from Nikita Popov worth 1.2 percent on its own, which for a single dependency bump is a large number.
This series is the closest thing the front page has to a standing instrument: same methodology, published regression count, no framing. Given the rest of today’s page, that is worth stating plainly. Nethercote has been measuring the same thing for years, reports the 74 benchmarks that got worse next to the 555 that got better, and the result is 233 points and no argument in the comments, which is the correct outcome for a claim backed by a reproducible measurement.
Git 3.0’s upcoming SHA-256 default will be a costly mistake — 231 points
231 points · 236 comments · blog.gitbutler.com · HN discussion
Scott Chacon argues that Git 3.0’s plan to make SHA-256 the default content hash is an expensive, valueless, avoidable global migration, and he has been sitting on the argument for years. His technical case is careful about what “broken” means: SHA-1 is broken in the sense that collisions can be manufactured on purpose for tens of thousands of dollars of GPU time, not in the sense that two reasonable files accidentally collide. The realistic threat is replacing a file’s contents with a malicious version that hashes identically.
Then he goes after the threat model, which is the part worth reading. A chosen-prefix collision is a terrible way to get untrusted code onto a machine compared with social-engineering a tired maintainer into handing over an unloved but widely depended-on npm package for a one-off payment — cheaper, faster, and undetectable in a way a collision is not. His second argument is about where trust actually lives: nobody pulls Rust from GitHub because they verified a commit SHA against a GPG signature and concluded that arbitrary sources are fine; they pull it because they trust that GitHub’s authentication holds and the maintainer did not push it. Given that, the proposed fix is not a global hash swap that bifurcates every tool, cache, signature and mirror in the ecosystem, but signing two hashes — keep SHA-1 for addressing content, add a SHA-256 (or BLAKE3) tree hash and sign both. The thread splits on whether that is elegant or merely defers the migration, and nobody disputes that the ecosystem-wide cost is real.
5x faster Edge Functions: V8 isolates to Firecracker MicroVMs — 226 points
226 points · 102 comments · netlify.com · HN discussion
Netlify runs about a billion edge functions a day and has rebuilt the infrastructure underneath them, moving from a hosted execution service to MicroVMs — Firecracker, working with Unikraft — inside its own edge network. The claim is roughly 5x faster at the median, with better security and reliability and no change to how functions are written or used. The blog post is co-written with the vendor whose VMM they adopted, which is worth noticing but not disqualifying.
Two caveats. The median is the flattering statistic for a latency distribution, and the reason V8 isolates were the default for edge compute in the first place was density — you can start thousands of them in the memory one VM would take, so MicroVMs win on cold-start consistency and lose on packing. Netlify is claiming the trade now pays, and the interesting missing numbers are p99 and the cost per invocation, which is what actually determines whether this is a V8-versus-VM result or a re-architecture result. The second caveat is that the same shift is happening everywhere at once: the edge platform that used to be a proxy is becoming a place to run real workloads, and this is one of three infrastructure posts on the front page today making that case.
Cloudflare K2: serverless event streams — 204 points
204 points · 83 comments · blog.cloudflare.com · HN discussion
K2 is in public beta: a durable event streaming primitive on Cloudflare’s developer platform. You write events to a stream, K2 stores them as an ordered log, and consumers read them either split across a set of readers for scale or delivered to all of them for fan-out. The stated problem is the usual one — RPC couples producers and consumers in time and scale, so when a consumer is slow or down, events are dropped, and multiple independent readers make it worse. The example is an ecommerce backend emitting transaction events that both an analytics system and a fraud detector need to see.
Kafka’s shape, with the operations replaced by a platform. That is a real product, and the interesting questions are the ones the post does not ask: retention and replay semantics beyond the phrase “long-term”, whether one region’s stream is readable from another, what a consumer group’s ordering guarantees are under partition, and how much work it is to leave. It is also the second Cloudflare post on today’s page, alongside Clef, and the pattern across the two — decision models, streams, and previously object storage, databases and VMs — is a company turning every primitive into a line item on the same bill.
FTC is investigating OpenAI, Anthropic and other AI companies over product risks — 200 points
200 points · 149 comments · cnbc.com · HN discussion
The FTC has opened an investigation into OpenAI, Anthropic and other AI companies over the potential dangers their products pose, confirmed by a spokesperson. The spokesperson declined to name the other companies, and the probe follows months of scrutiny of the two named labs’ safety practices. CNBC’s video segment frames it as a probe of “super intelligence models.”
An investigation is not a finding, and the story as published is confirmation that a probe exists plus the name of one agency. The substantive question underneath it is one nobody in the thread can answer either: what authority the FTC has over model behaviour, as distinct from advertising claims, data practices, or unfair acts — which are the things it normally regulates. If the theory of the case is that a model which produces harmful output is an unfair practice, that is a much larger claim than any prior consumer-protection action, and it will be litigated either way.
Most data centers refusing to say how much water, electricity they use — 216 points
216 points · 197 comments · nltimes.nl · HN discussion
A Lighthouse Report investigation with Trouw and other European newsrooms finds that the vast majority of European data centers keep their environmental footprint secret, and in the Netherlands fewer than a quarter of the larger ones publish anything about electricity and drinking-water consumption. The European Energy Efficiency Directive has required facilities with at least 500 kW of installed capacity to report energy and water use for three years. Enforcement has not followed. Dutch coverage from NOS, Trouw and NRC adds the hyperscaler angle, including reporting earlier this year that Google and Microsoft withhold hyperscale energy figures from a government that says it has no legal means to compel them. Both the NL Times page and Firecrawl hit a bot challenge, so the figures here come from the search citations and the Dutch-language coverage.
The thread is a small masterclass in how a disclosure argument fails. One side points out that environmental reporting requirements apply to companies that are not government agencies, so the interesting question is not why they withhold but why the directive has no teeth — three years of non-compliance with no penalty is a law-shaped suggestion. Another commenter argues the numbers are being withheld strategically: any published figure will be compared to something emotionally loaded rather than to agriculture or industry, so the rational move is to publish nothing and let the argument stay about feelings. The middle position is the one the story supports: water at market rate plus reporting as a condition of connection would settle it, and neither is happening.
Still on the page
Nine stories from the last four days are still above 200 points in today’s window, and the movement tells the usual story. You said no MCP 529 → 669, up 140 and today’s biggest mover, which is a reversal post having a very good week. Livenerf 853 → 911, up 58 — the model-nerf detector picked up more points than the model it was built to watch, which spent the day gated behind Fairwind. Solving Factorio Quality 240 → 274, up 34. America.gov 739 → 769 and Dots: Always-on agents 733 → 763, both up 30. Vermont replacing power plants with home batteries 357 → 383, up 26, NASA asked several former SR-71A staffers 300 → 320, up 20, Phyllotaxis 334 → 349, up 15. And When did Google get so weird? 1,987 → 2,002, up 15 — five days at the top, now 311 points clear of Gemini 4 Argon’s 1,643, and still gaining slowly. That is a strange shape for a story this old: it has not stalled, it has just become the page’s permanent resident.
Thirteen more from yesterday’s list fell out of the scanned window and were recovered through the item API, with every number verified against the story it actually belonged to: Updated Google Maps shows destruction of Rafah 922 → 936, Everybody’s home. No one’s coming over 807 → 819, Pirating the Pirates 695 → 701, It’s Time to Investigate the AI Labs 618 → 624, Ember-1 586 → 588, Coding is not solved 566 → 578, 500k facial scans at UK stations 504 → 512, Kids turned NPR Spotify comments into a group chat 478 → 483, A Privacy Analysis of Conversational AI Agents 421 → 423, MongoDB CEO resigns to join Meta 362 → 363, Does Reddit have an astroturfing problem? 303 → 309, World Labs joins AMD 305 → 307, Hijacking the PS5’s RTMP stream 290 → 292. Nine others from yesterday — including GPT 6.1 Sol at 1,039 and Jeff at 570 — could not be matched to a current score with any confidence, so they are omitted rather than estimated. The Chromebook support cut resurfaced under a new headline at 343 points, a resubmission of yesterday’s story rather than new coverage; Google’s change to ten years of updates is still the same eight years it was yesterday.
Throughline
First: the day was about the cost of state, not the size of the intelligence. The most consequential story on the page is a cache. DeepSeek’s KV compression reached 890 bytes per token — 437x below DeepSeek-V1, 99.77 percent less — and the visible consequence is commercial rather than technical: Anthropic cut Opus 5.5’s cache-read price by 60 percent and OpenAI cut GPT-6.1 Sol’s cached input by 80 percent, because holding context in VRAM was the largest serving cost for exactly the workloads everyone is selling. Everything else on the page rhymes with it. Argon’s headline number is 1M output tokens, which is a statement about how much expensive state a single trajectory can carry. Clef exists to make decisions bounded and cheap at $0.24 per million input tokens against a $0.042 incumbent. turbopuffer is rewriting its storage layer because maintaining the index that owned document placement has stopped paying. Netlify replaced isolates with MicroVMs to change a latency distribution. K2 sells the log between producers and consumers. Vendors cannot fake unit economics with a benchmark, which is why the only price movements in this roundup are downward and the only benchmark movements are upward.
Second: almost every claim today arrived with a caveat, and the ones without caveats were the weakest. Gemini 4 Argon leads the page with 1,643 points and cannot be run by anyone reading about it, with benchmarks nobody can reproduce. Magnet Forensics claims a capability in a leaked promotional video, with no CVE and no reproduction. A 16-year-old’s Microsoft write-up is more interesting for the disclosure that Microsoft edited it than for the token-signature bug. TLA+’s best-known educator used a viral result to point out that verification is limited by which properties someone thought to express. Livenerf’s most honest finding is the negative one. Micron’s supply warning is guidance from the seller. Meta’s tax position is legal until it is tested. The EU’s data center reporting directive is three years old and unenforced. Git’s hash migration argument is really an argument about where trust lives — in the repository source, not the digest. The page’s strongest entries were all instruments: a two-monthly compiler benchmark that reports its own regressions, a museum photograph of a panel that cannot lie to you, and a validation harness that says out loud what it cannot see.
Third: the fight has moved to distribution. Google will sell nobody the frontier model and will release it to defenders without guardrails first. Google’s Play console closes developer accounts for inactivity with no restoration path and no explanation, and the escape hatch out of the store — shipping an APK — now requires $25 and an identity verification, so sidestepping distribution costs the same compliance as joining it. Meta converts capex into a research credit and Europe’s data centers decline to report numbers the directive already requires. Meanwhile the most-distributed software on the page came from places with no gate at all: a KV-cache technique published by a lab with no incentive to hoard it, a bit-exact DLSS reimplementation that ships the code and asks you to bring the weights, a dial-up simulation that gets the latency right, and an iOS port of an OpenStreetMap editor paid for by a research ministry and a foundation. Yesterday’s page looked like a war over pricing metrics. Today’s looks like a fight over who is allowed to hand out the software — and the answer, in most of today’s stories, is not the platform.
Published 1 October 2026. Scores captured at 19:50 PDT; the front page moves after that.