Forty-two stories cleared 200 points today, seventeen of them new. The top of the page is unchanged for the fourth day: “When did Google get so weird?” added 40 points to reach 1,987 and is now 948 clear of the runner-up. That runner-up is a pricing announcement.
The new stories are unusually good. There’s a benchmark whose entire purpose is to detect whether a model you pay for got quietly worse, a government website that routes your passport question through two commercial LLMs, and a data-heavy essay arguing that the last thirty years were one long trade of buffers for dependencies. Two of the three are about verification.
Livenerf: Has Opus 5.5 been nerfed yet? — 853 points
853 points · 364 comments · github.com/ninjahawk/livenerf · HN discussion
For months people have claimed Anthropic “nerfs” models days or weeks after release — quantization, a smaller model behind the same name, lowered reasoning effort, routing changes. The counter-claim is that people are pattern-matching on noise. Nobody had a clean day-zero baseline, so the argument was vibes versus vibes. Opus 5.5 shipped on 2026-09-22, and this project started the clock on launch day. It runs once a day for 30 days: days 1–10 are the baseline, then two ten-day windows, so the first possible verdict lands around 2026-10-24. Today’s progress note says 7 of 30 days collected, none missed, all on the same harness hash and a pinned CLI version.
You can’t make these models deterministic — sampling parameters are gone and thinking can’t be turned off — so everything else is frozen: prompts, CLI, graders, and raw logs kept forever. It’s built on Inspect, the UK AI Security Institute’s eval framework, and the statistics follow Anthropic’s own paper on adding error bars to evals, which is a nice touch: the vendor wrote the method being used to audit the vendor. The panel is 78 questions scraped out of 2,336 GPQA Diamond, MMLU-Pro, competition-math and AIME items, selected for being the ones Opus 5.5 gets right only sometimes.
What makes this worth reading is that the author measured the instrument’s own failure modes and published them. Lower effort shows up far more clearly in tokens than in accuracy: effort at low gives −62 percent output tokens and −8.3 ± 4.5 points of accuracy, effort at medium gives −26 percent tokens and −4.2 ± 3.9 points. And the fatal caveat, stated plainly: swapping in plain Opus 5 was not distinguishable from Opus 5.5 at 99 percent confidence (−3.8 ± 6.3 points, −23 percent tokens). The instrument cannot detect a same-family model swap of that size. Which is the specific nerf people complain about most, and the one this benchmark cannot see.
Read the two numbers together and you get the actual result. If a lab quietly lowered reasoning effort, you would see your token consumption fall before you saw accuracy move. The instrument that detects a nerf is your billing dashboard, and the comment section converges on that independently — one commenter reports having depleted 54 percent of a Max allowance in two days this week after failing to deplete it in six days of sustained work on launch week. The thread also has a runner-up posting a second tracker (Nerf Bench, which pins launch-day behaviour and flags deviations above 10 percent, and claims to have caught the Opus 4.6 degradation Anthropic later blogged about), a skeptic arguing that people sense nerfs more often than they happen, and a rebuttal from someone pointing out that thousands of engineers ship 100,000-plus small changes a day, so some domains regress without anyone deciding anything.
America.gov — 739 points
739 points · 675 comments · america.gov · HN discussion
A single government front door: describe what you need in plain language and an AI answers using only official sources. The headline number is 29,000 federal websites collapsed into one interface, against the administration’s own framing that 40 million people a day interact with a government site trying to get something done. It’s free, has no ads, keeps no account, and the landing page states that your conversation disappears when you leave. Phase one is answers. Phase two, starting in 2027, is transactions — apply, enrol, and track progress inside the chat.
The chatbot is the least interesting part of the announcement, and the executive order signed at the launch is the rest. It directs agencies to integrate public-facing services into America.gov within 90 days, covering services that serve more than 100,000 users over 12 months and can be applied for online — while explicitly excluding IRS tax filing and national-security services from the Department of War and the intelligence community. That carve-out is the story. The tax filing system is the single most painful annual interaction most Americans have with the federal government, and it is out of scope by decree. The one front door excludes the room most people are trying to get into.
Two things are worth being precise about. First, the models: U.S. Chief Design Officer Joe Gebbia told CNBC on launch day that the site is powered by Google’s Gemini and Elon Musk’s Grok, which is a specific claim that neither the site nor the executive order makes. It was built by the National Design Studio — an August 2025 executive order creation staffed in part by former DOGE employees, led by the Airbnb co-founder — and is operated by the GSA. Second, the privacy claim. “Your personal information isn’t collected or stored” is a statement about the application layer, and it sits above a pipeline that hands your question to two commercial model providers and an agency backend. I tried to read america.gov’s own privacy and how-it-works pages for the specifics and both sit behind a Cloudflare bot challenge, so I could not verify what the policy actually says. The animated fingerprint that sweeps across the words “keeps your privacy protected” does not dismiss when you click it.
The comment section is more useful than usual on this one. The top comment is enthusiastic and correct about the actual problem being solved — people getting phished because they can’t tell which of 29,000 sites is real — and someone found Google’s own blog post confirming the Gemini partnership and a “100 million people” reach claim. Then there’s the xkcd-shaped joke that lands because it’s the right criticism of every consolidation project: there are 29,000 competing government websites. Ridiculous! We need to develop one government website that covers everyone’s use cases. One commenter notes the irony of shipping a reinvention of government services behind a bot-detection wall.
You said no MCP — 529 points
529 points · 311 comments · earendil.com · HN discussion
Pi’s makers spent a year publicly dismissing the Model Context Protocol — the landing page said the tool didn’t support it, podcasts said worse, one of the founders wrote a piece titled “what if you don’t need MCP.” It’s now in core, and this post is the explanation. The short version is that MCP changed, but that alone wouldn’t have justified core support, because MCP could already run as an extension and did.
The real reason is in the middle of the post and it isn’t the one the framing advertises. The changes they needed for MCP also enable running Jev inside Pi, because both need the same thing: a sandbox to play in, in the form of an interpreter. MCP in Pi is built on exposing tools to a JavaScript sandbox, the same approach Codex takes. They didn’t reverse a judgement about whether MCP is good — they found that supporting it cost them less than they’d assumed, because they were building the interpreter anyway. That’s a better argument than “the world is not static,” and it’s the honest one.
What hasn’t improved, by their own account: MCP is still hard to compose. Codemode — “a neat little sandbox to allow composing of tool calls” — doesn’t fully solve it, and they place the blame on how servers are built rather than on the protocol. Many still target harnesses that dump every tool into context and optimize on their side by returning text. Their preferred shape is MCP as something closer to OpenAPI with intelligent tool discovery: tools returning structured data, discoverable by their own documentation. That’s the part of the post worth arguing with. Structured returns are better, but “discoverable by documentation” is doing a lot of work when the documentation is written by the server author and read by a model.
The thread’s best exchange is about the reversal itself rather than the protocol. One commenter praises the team for changing their minds in public and links Armin Ronacher’s 2016 piece about being careful what you dislike, quoting the line that arguments about a strong opinion are usually fought with reasons that expired years ago. Underneath that sits a much larger MCP use case than coding: several macOS apps now ship MCP servers so users can configure them in natural language — “set up Clop to optimize any PNG I drop in my website assets folder.” That’s a genuinely different argument for MCP’s existence, and it isn’t the one the post makes. The best skeptic line is the compatibility one: it’s suboptimal for the reasons the author outlines, but so are USB-C, NVMe and HDMI.
September 2026: The world today, as seen by one Polish guy — 472 points
472 points · 357 comments · tomwojcik.com · HN discussion
A Polish writer’s heavily charted read of the year, built on one thesis: for thirty years we swapped buffers for dependencies, because a supplier is cheaper than a stockpile and a guarantee is cheaper than an army. When a dependency failed, we didn’t rebuild the buffer — we found another dependency. Every swap worked as long as the thing at the other end was still there.
The counting is the strong part. US and Israeli operations against Iran began in late February; since March, Iran has kept the Strait of Hormuz closed with drones, missiles, mines and small boats. Tanker traffic through it is down more than 90 percent, and the IEA calls it the largest supply disruption the oil market has ever seen. A fragile ceasefire pulled prices back to pre-war levels in early summer, then broke down: Brent near $97 in early September, up 19 percent in a month, about $105 by mid-month, $108 on the 24th. On 22 September Iran handed Washington a written road map — a regional ceasefire of up to 60 days, phased reopening of the strait, an end to the naval blockade — and Washington rejected it. The detour around the Gulf runs through Bab al-Mandab, where Houthi forces seized a key Yemeni port this month. And the same closure that empties a granary fills a brokerage account: the Breakwave Tanker Shipping ETF rose more than 600 percent in the first two months of the war and more than 2,300 percent for the year by early September, with supertanker day rates going from under $100,000 before the war to a record $860,000 on 10 September.
The essay’s best single line is the one about reversibility: a Qatari plant needing three to five years of repairs does not care about the next ceasefire, and fertiliser that wasn’t spread this season can’t be spread in retrospect. Relief is not repair. The ceasefire already happened once this year and reversed, and nothing about a reopening would have restored the buffers.
Where it’s soft is where it aggregates. The fertility-to-robots chain is a forecast and the essay says so in passing — Poland’s fertility rate hit a record low in 2025, population decline accelerated with 168,000 more deaths than births, the over-retirement-age share went from 12.8 percent in 1990 to 24.2 percent in 2025, and the answer on offer is machines. More than half the factory robots installed worldwide in 2024 went into Chinese plants; Chinese suppliers’ share of their own home market climbed from 30 to 57 percent in four years; by one industry estimate Chinese firms account for close to 90 percent of humanoid robots shipped globally. The robotics federation’s own caveat, which the essay quotes, is that humanoids in real production are still demonstrators and pilot projects. So the conclusion — that a continent replacing missing workers with machines made in China, running software written in America or China, is taking on a third dependency on top of energy and defence — is a hope dressed as a plan, and he admits it.
The Polish commenters push back hardest and the pushback is the most valuable thing in the thread. One calls the piece overdramatic and points out the country’s storage is at 97 percent because a previous administration was paranoid enough to build a marine gas terminal and the current one stocked up early; another notes inflation-adjusted property prices peaked two and a half years ago. That is an argument for the thesis rather than against it — the buffer was buildable, it just required someone to be paranoid in public before the crisis. The essay’s own conclusion is deliberately modest: prepare for the boring version, hours without power rather than months, and keep more cash than feels clever.
Ask HN: What are you reading? — 407 points
407 points · 812 comments · HN discussion
The recurring thread, and today it drew the second-highest comment count on the page. The list is more literary than this site’s reputation suggests. Donella Meadows’ Thinking in Systems — given as a birthday present by a 17-year-old son, and praised for putting names to patterns the reader already understood. Two Vernor Vinge novels, A Fire Upon the Deep and A Deepness in the Sky. Larry McMurtry’s Lonesome Dove, described in the best framing of the thread as “horsepunk: instead of cyberpunk’s high tech low-lifes, it’s basically horseback low-lifes.” A book by an author named Dienw Neb titled null: $cat /dev/null_, discovered in an ad in the back of 2600 Magazine, whose IRC chat recreations the reader says are too accurate to be invented. Plus Ted Chiang, Ken Liu’s The Paper Menagerie, Cryptonomicon, Stoner, East of Eden, The Count of Monte Cristo, and a confession from someone who went straight from the Three-Body trilogy to Dungeon Crawler Carl and describes the drop as “feels like cheating, but it’s making me happy.”
The ratio is the interesting part: 812 comments against 407 points, roughly two comments per point where the page’s median story runs closer to half that. A recommendation thread is one of the few formats on this site where everyone has standing, no one can be wrong, and the only cost of contributing is admitting what you actually like. It’s also the closest thing HN has to a census.
Show HN: Real-time Solar System with 526k asteroids and all tracked satellites — 371 points
371 points · 102 comments · space.bl2.net · HN discussion
A WebGL2 browser view of the solar system at real scale in its current state, plus everything catalogued around Earth. Data is CelesTrak TLEs propagated with SGP4, asteroids and comets from JPL’s small-body database, spacecraft positions from JPL Horizons, updated daily. Orbit propagation runs in web workers; the asteroid set is about 30 MB and streams in the background. The time slider runs backwards as well as forwards, and satellites appear and disappear by launch date, so you can scrub to any point and watch the constellation build.
The census as of this morning: 19,890 satellites with published orbits plus 15,252 approximate, 837 solar-system bodies, and a debris breakdown that is the actual point. Fengyun-1C debris from the 2007 Chinese ASAT test: 1,981 tracked objects. Cosmos 2251 from the 2009 collision: 583. Iridium 33 from the same collision: 111. Cosmos 1408 from the 2021 Russian ASAT test: three. Then roughly 9,681 other debris, 2,233 rocket bodies and 3,274 inactive satellites held at approximate positions. Four separate events, fifteen to nineteen years apart, each still contributing hundreds of individually catalogued fragments.
The best thing on the page is a disclaimer most visualisations of this kind omit: distances are to scale but dots are not, because a dot is about 81 km across while a satellite is a few metres. Every picture of orbital congestion is lying about collision probability by roughly four orders of magnitude, and this one says so out loud. The other number worth holding is the Starlink line — 11,127 satellites in the “active” bucket against a debris field of a few thousand, which is the shape of the problem: it isn’t the junk that’s growing fastest. One commenter notes having run orbital calculations for 30 objects on a 486DX and getting 30 frames per second; another observes that unchecking the satellite layer turns the whole scene placid.
A Staff Engineer’s Guide to Inventing Work — 359 points
359 points · 79 comments · sujithjay.com · HN discussion
Platform teams are engineering-led rather than product-led: there is no product manager handing over a roadmap, no revenue line to follow, and no market to lose. So work does not exist unless an engineer invents it, and a staff engineer’s job is largely to invent it. The guide reads four sources of signal — the system, the users, the organization, the industry.
Two of the signal types are worth naming. Crash-led discovery: postmortems tell you what to fix or replace, and only occasionally what to build, but the drawback is stated honestly — it biases the team toward the loudest and most recent failures rather than the largest opportunity, and it is a maximally lagging indicator. Cost: the cloud bill substitutes for a revenue north star, but the advice is to go past query tuning into your business unit’s P&L, vendor contracts and cost-centre line items, and to ask for each one whether the function should be pulled in-house or pushed out.
The comment section finds the flaw the framing licenses, and finds it in the first comment. A platform team with no customer cannot be wrong, so “inventing work” becomes self-justifying: “it’s precisely because of this framing and mentality that platform teams don’t actually serve people well and are usually highly dysfunctional towers of people inventing work.” The proposed fifth signal is the only one that costs something — act as if the teams you serve could leave, because they usually can. Another commenter, a staff-plus platform engineer, says the diagnosis is backwards: it’s obvious what increases user value, and what’s genuinely hard is building a case for why anything should be worked on at all when your team sits furthest from the customer. A third is blunter and right: if you can’t reason about your work in business metrics, you’re doing an academic exercise.
Vermont replacing power plants with home batteries — 357 points
357 points · 265 comments · bbc.com · HN discussion
Green Mountain Power leases participants two home batteries for $55 a month on a ten-year agreement. More than 5,500 Vermont homes are enrolled; one participant is on record comparing the lease favourably to the $12,000 gas generator she was about to buy. GMP serves over 75 percent of the state and has a target of eliminating all outages by 2030, and its aggregated home batteries have quietly become the state’s largest source of power. Seven of the ten most damaging storms in the utility’s history have hit in the last decade, causing more than $225 million in damage.
Zoom out and the programme is a small piece of a fast-growing category. The US has over 40 GW of virtual power plant capacity according to Wood Mackenzie — small against roughly 600 GW of natural gas and 200 GW of coal, but a 2025 Department of Energy report puts 160 GW as unlockable by 2030, equal to about 20 percent of expected peak demand.
Two corrections to the article’s framing, both from the thread and both fair. First, “Vermont’s largest power source” is a capacity claim, not an energy claim, and a battery is a consumer of electricity rather than a producer of it: the value is moving generation from off-peak to peak, minus round-trip losses and degradation. Second, the savings mostly come from avoiding peak demand charges rather than from energy arbitrage, which is a billing artifact as much as an engineering one. The financial structure deserves the skepticism it gets — the ratepayer buys hardware on a ten-year lease that the utility then dispatches, and one commenter calls that shifting utility costs onto consumers while retaining the right to use all the stored energy. The counter is factual: GMP has already retired two peaker plants because of the programme. Both things are true, and the honest reading is that Vermonters are buying the utility’s peaking capacity and receiving resilience in exchange. Whether that’s a good trade depends on what a peaker plant would have cost them instead.
PS5 Relapse Exploit — 343 points
343 points · 232 comments · github.com/ntfargo/Relapse-Exploit · HN discussion
A jailbreak chain covering PS5 firmware 7.00 through 13.60. The browser stage uses JavaScriptCore information leaks and a structured-clone object-pool mismatch to corrupt a typed array; the kernel stage combines an address leak with an aio_multi_wait use-after-free race to establish kernel read/write. Payloads land in payloads/ after a successful run and an ELF loader listens on port 9021. The README is unusually candid about reliability: WebKit may need several attempts, the kernel exploit may hang or panic the console, and you should reboot before retrying.
The firmware range is the number to sit with. 7.00 through 13.60 is roughly four years of system software, and the load-bearing kernel bug — an async I/O wait race — went unpatched across all of it. That says less about Sony’s patch cadence than about the maintenance surface of a console OS: the components with the worst bug density tend to be the ones nobody is looking at. Credits are split cleanly between the kernel author, the WebKit author, and the chain builders, which is how this scene has worked since it had different targets. The educational-purposes disclaimer is boilerplate, and the countermeasure is obvious — pin the payload hashes and validate the initial entry point.
NASA asked several former SR-71A staffers to help secret restart — 300 points
300 points · 344 comments · aviationweek.com · HN discussion
Mike Relja, a retiree in his late seventies who was a US Air Force and NASA test engineer on the SR-71A, got a call about a month ago asking for his help returning a Blackbird to flight after a nearly 27-year hiatus. His answer is the most quotable thing in the piece: “I wish you well. I’d like to see you make it. It would be interesting. But I just don’t think you can get there from here.” The aircraft is Tail No. 844, removed earlier this year from a display pole at the Armstrong Flight Research Center at Edwards, where it had stood since about 9 October 1999 — the day of the last flight of 844 and of any Blackbird. Satellite imagery obtained by Aviation Week dates the removal to before 17 May and perhaps before April. The request came from David Ash, an ex-Navy and ex-Joby test pilot hired in June as a special projects director, and the agency administrator is publicly hunting new X-planes.
The airframe is the least interesting part of the story. A man in his late seventies is the state of the art for a 1960s titanium aircraft because the industrial base that produced it is largely gone — the jigs and tooling destroyed in the 2000s, the titanium supply chain, and the people. Relja’s assessment is the engineering one and Ash’s hire is the management one, and they point in opposite directions. The thread is right that this is probably a budget play before it’s an engineering programme: the money goes to national prestige projects and things that feed hypersonics, and a retro-flyable X-plane is a cheap way into that conversation. Underneath the cynicism there’s a real argument, made better by one commenter than by the article — that we let whole areas of engineering die without attempting to preserve them, and we discover the gap only when we need the capability back. The counter is that re-flying a sixty-year-old airframe is the most expensive possible way to preserve expertise. Note also that the article is free to read until 29 October.
Backblaze drive stats for Q2 2026 — 295 points
295 points · 102 comments · backblaze.com · HN discussion
Thirteen years of publishing failure rates for a fleet nobody else will describe in public. Q2 2026 ran April 1 to June 30: 359,101 drives monitored, 3,881 boot drives and 705 HDDs excluded, leaving 354,415 drives, 31,553,350 drive days and 1,498 failures. That’s an annualized failure rate of 1.73 percent — Backblaze’s own words are “the highest it’s been in quite a while.” Lifetime AFR through June 30 is 1.41 percent across 353,041 drives and 558 million drive days. Three outliers came in above the 6.95 percent quartile threshold: an HGST 12 TB at 7.63 percent, a Seagate 10 TB at 9.33 percent and a Seagate 14 TB at 8.26 percent. Backblaze gives the reasons rather than burying them — the HGST unit is nearly seven years old, the 10 TB Seagate nearly eight and a half, and both Seagate models have small populations, so 22 and 25 failures produce a very high rate. A Seagate 8, 12 and 14 TB model each posted zero failures and a 16 TB posted one, which is the clean sweep in the post. Two drives retired at nearly nine and nearly eight years old, and for the second consecutive quarter there were no new drive models at all, with deployments concentrated in 20 TB and above.
The headline needs splitting, because it’s a mix effect as much as a reliability signal. The drives dragging the average up have average ages between 59 and 115 months; the models absorbing new deployment sit at 3 to 19 months with AFRs of 0.55 to 1.16 percent. An aging fleet getting older produces a rising aggregate AFR without any drive getting worse.
The best thing in the thread is a decade-long series that argues the opposite of the headline. In 2013, Backblaze’s fleet hit roughly 14 percent failure at three years three months, implying a useful life near four years. By 2021, 14 percent arrived at seven years nine months. By 2025, about 5 percent at ten years three months. The three-year failure rate that was 14 percent in 2013 is now well under 5 percent, and that is a much more important chart than this quarter’s 1.73 percent. The other recurring note is capacity without throughput — drive sizes keep climbing while sequential speeds haven’t, so reading one of these end to end takes days, which makes them poor RAID rebuild candidates and worse cold storage than the marketing implies. Backblaze’s own caveat is the honest one: capacity and AFR matter, but so do sustained write behaviour, rebuild performance, parity overhead and firmware — and with SMR drives entering the fleet, the AFR stops being a single-vendor number.
Tcl/Tk 9.1 — 293 points
293 points · 139 comments · tcl-lang.org · HN discussion
Tcl/Tk 9.1.0 shipped on 29 September. The Tcl side adds a unicode command for normalization, a timer command with a monotonic clock at microsecond resolution, lfilter, interp set for child-interpreter variable access, -backslashes/-commands/-variables options to subst, -integer for switch, C99 math routines as expr functions, and a new C time API using long long in place of Tcl_Time. Many list and hash helpers gained C routines, list internals got more memory-efficient for large lists, and 64-bit size support was extended. Tk picks up accessibility screen-reader support, initial bidirectional and right-to-left text support, a ttk::toggleswitch widget, and tk attribtable.
Two of those entries matter to anyone outside the community. Screen-reader support in Tk is a decade-late accessibility fix for a toolkit that still ships inside commercial products — and for the audience Tk actually serves, the long tail of internal tools and lab and industrial software, it may be the most consequential line in the release. The other is a breaking change: applications are now required to call an initialization routine, either Tcl_FindExecutable or TclZipfs_AppHook. That will surface as a crash on startup for embedded users who skim release notes.
The thread’s best contribution is a historical footnote that reframes the whole language. D. Richard Hipp, quoted in the comments, has said that SQLite is “a TCL extension that has escaped into the wild,” that SQLite’s design was inspired by Tcl in both its datatype handling and its source-code formatting, and that the founding use case was a Tcl/Tk application at an industrial company. Tcl is one of very few 1980s scripting languages whose design decisions run on essentially every phone and browser — just not in the language itself. The rest of the thread is affection rather than argument: everyone’s easiest GUI toolkit, everything is a string, and one commenter’s line about multi-interpreter threading being something Tcl got right before anyone else did.
1 in 8 cancer cases worldwide are caused by infections, study finds — 283 points
283 points · 148 comments · cbc.ca · HN discussion
Researchers at the International Agency for Research on Cancer, a WHO arm, published Monday in The Lancet Oncology: an estimated 2.3 million new cancer cases in 2024 — 12 percent of the total — are attributable to infectious agents. The largest contributors are Helicobacter pylori at 760,000 cases, concentrated in eastern Asia, and HPV at nearly 750,000, concentrated in sub-Saharan Africa. HPV is responsible for 100 percent of cervical cancers plus anal, vulvar, vaginal, penile and some head and neck cancers. Hepatitis B and C account for a substantial share of liver cancer. For the first time the analysis extends the list of infection-linked cancers to include 16 more.
Be careful with the causality in the headline. These are population attributable fractions — an accounting estimate of how much disease would not occur if the agent were absent — not a mechanism. That distinction matters most at the margin, and the margin is where this study grew: expanding the attributable set to 16 additional cancers is what moved the number, and cancers joined with weaker associations raise the fraction without raising confidence in each link. The thread’s best question is the confounding one, and it is not answerable from this data: if an immune system that fails to clear an infection is also the immune system that fails to kill a nascent tumour, the infection might be a marker rather than a cause. Worth remembering that viral infection was the dominant theory of cancer until the 1970s, before the field pivoted to genetic drivers — the pathogens never went away, they just stopped being the story.
Where the finding is unambiguously actionable is the top of the list. H. pylori is treatable with a course of antibiotics and HPV has a vaccine, and both are concentrated in exactly the regions with the least screening and the lowest vaccination coverage. A 12 percent figure that is dominated by two preventable or treatable agents is mostly an indictment of distribution, not of biology.
U.S. postal inspectors shut down website selling counterfeit postage labels — 273 points
273 points · 163 comments · postalemployeenetwork.com · HN discussion
The Postal Inspection Service shut down a site selling millions of counterfeit postage labels; the operator is described as Pakistani and has been charged, with no reporting that he’s been arrested.
The article is thin and the comments carry it. One commenter is fairly sure they bought from a seller using this service on eBay, and walked the chain: tracking number issued for the right town, package that never arrives in the expected form, and a feedback profile sitting at “100 percent positive” despite enormous volume — so much volume that a thousand neutral ratings barely registers. That’s the actual finding. Marketplace reputation systems measure the transaction from the buyer’s side, and the buyer of a counterfeit-label shipment has no way to know the label was fake and every incentive to rate it well, because the goods arrived. The fraud is invisible at exactly the layer where the trust signal is generated.
There’s also a good digression from someone with direct experience of the enforcement side: USPIS visiting a federal prison to explain that stamps are currency and using them as such is a federal offence, along with a breakdown of “flats” versus “books” and their street value — around $13 to $14 for a sheet of 20 Forever stamps. Which raises the relevant asymmetry: counterfeit postage is a target for a federal taskforce while buying it is effectively unenforced, because the loss lands on a self-funded agency rather than on any identifiable victim.
Solving Factorio Quality — 240 points
240 points · 91 comments · exyr.org · HN discussion
The author plays Factorio the normal way, by writing matrix math to plan the factory. Specifically, he models quality — the five-tier system added by the Space Age expansion in 2024 — as a linear program. Quality can only be increased through quality modules, each tier step after the first is another 10 percent chance, the maximum quality chance in a four-module machine is 24.8 percent, and a single-step jump from normal to legendary is a 0.0248 percent event. From there it’s washing versus upcycling, the marginal value of a quality module against a speed module, and a solver. There’s a hosted calculator at the end of the post.
This genre keeps producing good writing because the game’s crafting graph is genuinely a linear program and its constraints are stated in numbers. Two comments are worth more than the article’s charts. The first corrects a claim in the post: Factorio did not found the factory genre, it was heavily influenced by Minecraft mods like IndustrialCraft2, and the creator was on that forum in 2013 saying so — the genre’s lineage runs Minecraft mod to Factorio, not Factorio ex nihilo. The second is a result the author doesn’t report: adding just a small number of speed modules goes a long way, especially at lower quality tiers, and the benefit of better quality modules is non-linear, so players systematically underestimate it. Also worth flagging that a patch has since removed the “Space Casino” the post relies on, which is the standard problem with any optimisation write-up for a live game: the numbers describe a version that no longer exists.
ChatGPT Pro 500 — 216 points
216 points · 258 comments · help.openai.com · HN discussion
OpenAI now sells three tiers of Pro: Pro 100 at $100 a month, Pro 200 at $200, and Pro 500 at $500, which is the only one that includes Astra Ultrafast. Pro 200 is available to new subscribers again, with the telling qualifier that subscriptions not eligible for grandfathering “include a lower usage allowance than previously offered to reflect our increasingly efficient models.” Existing Pro 200 subscribers who were active at the eligibility cutoff keep their old allowance through 29 October, after which they move to the reduced allowance at the same $200 price. Keeping the allowance does not add Ultrafast.
The thread assembled the surrounding facts faster than the documentation did. One commenter built the multiplier table — Plus at 1× for $20, Pro 100 at 5× for $100, Pro 200 at 10× for $200 where it used to be 20×, Pro 500 at 25× for $500 — and drew the correct conclusion: the scaling is now linear, twenty-five Pro subscriptions’ worth of usage for twenty-five times the money, and the “deal” in the old Pro 200 tier has been withdrawn. Another commenter went looking for the actual limits on all three pages OpenAI publishes and could not find them anywhere, including the signed-in upgrade screen. And a third quotes the launch email, which included a one-time grant of 62,500 usage credits valued at $2,500 and expiring 31 December, with no published definition of what a credit is. Four cents per credit is an implied price for an undefined unit of work.
Put this next to the Livenerf entry at the top of the page and the two stories are about the same thing. The nerf complaint and the pricing change are the same complaint: the product you subscribed to is being re-specified underneath you, and the only instrument that reliably detects it is the one that counts your tokens. A vendor cutting an allowance and calling it efficiency, then selling the difference back at a 2.5× price premium, is a strategy that works exactly as long as nobody can measure the thing they’re buying. Today’s front page contains two projects built specifically to measure it.
Google ending ChromeOS support two years early — 211 points
211 points · 141 comments · theregister.com · HN discussion
Google’s policy promises ten years of updates for Chromebooks. A support document titled “What the Googlebook announcement means for your ChromeOS devices” says that a Chromebook bought today will get updates until 2034 — eight years. The mitigation offered is transition support: “for qualifying devices purchased today whose 10-year support lifecycle extends beyond 2034, Google is committed to supporting your transition to Googlebook OS.” Googlebooks are the replacement line, with more powerful processors and a new Googlebook OS that bakes in Gemini.
The arithmetic error is the hook, but the policy is the story, and it’s worse than a typo. Ten years of support was the one durable thing Chromebooks offered schools — the reason a district could buy a fleet and amortise it. Truncating every new device to the same 2034 end date means a Chromebook bought in 2026 dies at the same moment as one bought in 2028, which is exactly the synchronised-replacement problem districts adopted Chromebooks to escape. And the migration path runs to an OS that isn’t finished, on hardware the education channel can’t currently buy: none of Google’s five Googlebook partners has built a model for education or other fleet buyers. The five, and many other OEMs, are still shipping Chromebooks.
The thread gets to the right place quickly. This is not retroactive — existing Chromebooks keep their ten years, and one commenter notes that plenty of eight-year-old machines are still in daily use. The line that will age best is the plain one: file this as another datapoint for evaluating Google’s long-term commitments, no, we mean it this time. There’s also a genuinely sharp observation in the replies about the timing — if the company’s own models are as productivity-multiplying as advertised, maintaining an OS for two more years should be cheaper than it used to be, not more expensive. Big software’s maintenance behaviour is showing no sign of being improved by the tools it’s selling.
Still on the page
Twenty-five of today’s forty-two stories above 200 points were covered here in the last four days, and the re-cuts moved hard today. GPT 6.1 Sol 562 → 1,039, up 477 and now clearly second on the page — yesterday’s price-cut framing has been absorbed and the story is now simply the leading OpenAI item. Dots: Always-on agents 353 → 733, up 380, the second-biggest move on the page, and between them the two DevDay announcements added 857 points in a day. How Delhi cut electricity loss 368 → 573, up 205, the largest jump by a story that isn’t about AI. Everybody’s home. No one’s coming over 632 → 807, up 175. Phyllotaxis 234 → 334, up 100 and still the quietest story on the page at 50 comments. Updated Google Maps shows destruction of Rafah 864 → 922 up 58, 500k facial scans at UK stations 449 → 504 up 55, and When did Google get so weird? 1,947 → 1,987 up only 40 — four days at the top and finally out of steam, which is the normal decay curve for a piece that got 1,100 comments. Then It’s Time to Investigate the AI Labs 587 → 618 up 31, A Privacy Analysis of Conversational AI Agents 390 → 421 up 31, Coding is not solved 538 → 566 up 28, Does Reddit have an astroturfing problem? 277 → 303 up 26, California farmers and wine grapes 363 → 383 up 20, Jeeves 206 → 239 up 33, Pirating the Pirates 681 → 695 up 14, Kids turned NPR Spotify comments into a group chat 465 → 478 up 13, Jeff 559 → 570 up 11. Grinding: MicroLLM Lab 272 → 280, Nvidia’s agent watchdog chip 219 → 226, HN.watch 207 → 213, MongoDB CEO resigns to join Meta 357 → 362, World Labs joins AMD 300 → 305, Hijacking the PS5’s RTMP stream 287 → 290, Ember-1 584 → 586, Parley 323 → 324. The bottom of the list is where the story is: five of the last six repeats moved by five points or fewer, and Ember-1 and Parley moved by two and one. That is what a front page looks like when the day’s attention has moved entirely to new material.
Throughline
First: the meter is the product now, and nobody will publish its units. GPT 6.1 Sol at a fifth of the price, ChatGPT Pro 500 at $500 a month with no usage table published on any of the three pages that describe it, Pro 200’s allowance halved at the same price and labelled efficiency, Dots to be scaled “by speed or monthly work volume,” and a one-time grant of 62,500 credits priced at four cents each with no definition of a credit. When capability stops being the scarce input, what’s left to sell is the measurement of consumption — and the vendors have every reason to leave the unit undefined. Livenerf exists precisely because of that gap, and its most honest result is a negative one: it cannot detect a same-family model swap at a validation’s worth of samples. The metric that would catch a quiet downgrade is a token counter, which is a number the vendor also controls.
Second: the day’s real output was verification instruments. Livenerf pre-registers a hypothesis, freezes a panel, and states that it can’t see the failure mode it was built to look for; Nerf Bench pins launch-day behaviour and reports deviations; Backblaze has been publishing the same numbers for thirteen years and this quarter explained its three worst drives instead of hiding them behind an aggregate; the solar system visualisation prints the disclaimer that its dots are 81 km wide and therefore useless as a collision-risk picture; Factorio’s quality math is a linear solver anyone can re-run; and a Polish essay’s charts are all sourced and all linked. Set against that: America.gov asserting your conversation is never stored and your privacy protected, with the privacy page behind a Cloudflare challenge and two commercial LLMs in the path. A page full of instruments to check claims, and one very large claim with no instrument attached.
Third: buffers versus dependencies, and the front page took both sides today. The Polish essay is the explicit version — thirty years of trading stockpiles for suppliers and guarantees for allies, described as a single design flaw viewed from a country with less slack than most, and answered in the comments by the observation that Poland’s own 97 percent gas storage was bought by an administration paranoid enough to build the terminal. Vermont is the literal version: 5,500 homes holding batteries, two peaker plants already retired, a utility aiming to eliminate outages by 2030, paid for by ratepayers leasing hardware they don’t control. NASA is the version you can’t buy back: a retired engineer in his seventies as the state of the art for a titanium airframe, because the tooling was destroyed and the people are the only remaining inventory. And ChromeOS cutting ten years to eight is the version being withdrawn: the fleet-wide durability that made Chromebooks attractive to schools, removed to synchronise the upgrade cycle onto a successor that doesn’t exist yet. Dependencies are cheaper right up until the moment they aren’t, and the invoice always arrives with less warning than the decision did.