Today’s front page holds thirty stories. Fourteen of them are above 200 points. The full 500-item rank list holds 124 — fifty-two inside rank 160, seventy-two more sitting in the long tail between rank 161 and rank 491 where the front page will never show them. Thirty-six of those 124 have no coverage in the last five roundups.
The page’s shape today is easy to describe: two announcements of frontier-scale open-weight models, neither of which has weights. Mistral says a trillion-parameter model is in public preview and the weights land by month’s end. Reflection says a 501-billion-parameter model is in final red-teaming and the weights land later this month. Together they hold 1,787 points, and neither artifact exists outside a lab. Aleph Alpha, with a 78-billion-parameter model that is already downloadable, holds 699 at rank 141, where the front page does not show it.
And the single biggest mover in the window is invisible. Anthropic reported diary entry to police gained 480 points — 315 to 795, with 630 comments — and sits at rank 36. The front page reported a placeholder domain’s CSS animation today instead. That is not a complaint about the voters; it is what rank-as-relevance produces.
Mistral Large 4 — 1,261 points
1,261 points · 811 comments · mistral.ai · HN discussion
Mistral launched a public preview of Mistral Large 4 — officially “le Chonk” — a one-trillion-parameter natively multimodal model with 49 billion active parameters. The weights drop at the end of the month. It was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral’s own European datacenters and served from the same hardware. A significant share of the training data was multilingual, spanning more than 160 languages including every official language of the EU.
The marketing is wrapped in sovereignty, and for once the wrapper is load-bearing. Mistral’s argument is that provider-level refusals block legitimate vulnerability research and incident response, and that losing access to a capability mid-incident is itself a security risk — so the pitch is a cyber-capable model with open weights and self-deployment. To that end, before release, it is being red-teamed by cybersecurity firms, “vetted partners” and state authorities with reduced moderation and expanded cyber capabilities. Read that twice: outside parties get an unlocked, cyber-specialized trillion-parameter model weeks before the public gets the locked one. It is the most defensible version of the argument and still the most interesting risk on today’s page.
The thread is 811 comments of everything HN does. Simon Willison notes the reasoning control has only two settings, “none” and “high”, and that high produced fewer output tokens than none — plus the pelican test, which is now the de facto vision benchmark in these threads. A commenter replies that the pelican is pristine and declares AGI arrived. The good argument is about capital, not capability: one commenter says the greatest threat to Mistral is that Europe has no capital market deep enough to absorb the valuation step-ups its sovereign growth investors need, and that this stays a tomorrow problem only as long as the company never goes public. An EU-based commenter makes the more useful admission — they will use US suppliers for coding, but never for functions inside their own products. That is what sovereignty demand actually looks like operationally, and it is a smaller number than the announcement implies.
Beam: Reflection’s 501B open-weight model — 526 points
526 points · 164 comments · reflection.ai · HN discussion
Reflection’s Beam is a sparse mixture-of-experts model, 501 billion total parameters and 23 billion active, pretrained on 23.8 trillion tokens from the web and “proprietary licensed datasets”. The number worth extracting from the press release is the RL run: over 100 million rollouts on 10,500 Nvidia GB300 GPUs across four weeks. That is the actual moat disclosure — not the parameter count, which is a table entry, but the scale of the post-training compute, which is a capital requirement almost nobody outside five companies and two governments can meet. Weights, technical report and model card come “later this month”, after final red-teaming.
Benchmarks put Beam at 44.4 on DeepSWE v1.1, 80.1 on Terminal Bench v2.1, 80.9 on SWE-bench Verified, 97.8 on AIME 2026 and 36.2 on HLE without tools. Reflection’s own framing is honest in a way these posts rarely are: it is competitive with GLM 5.2 and approaching Qwen 3.8 Max, Kimi K3 is ahead on raw capability, and Beam’s claimed advantage is inference-time efficiency. When a vendor tells you it is second on capability and first on cost, believe the first half.
The thread does the two things HN does to every pre-release announcement. First it audits the demonstration: one commenter catches Beam’s “generalization experiment” claiming a puzzle is “a few days old, so could not appear in the training data” — other commenters remember the puzzle from earlier, and the point survives only in the weaker form that they didn’t RL on it. Second it audits the entity: a commenter asks what kind of organization Reflection is, and the answer is Misha Laskin (reward modeling on DeepMind’s Gemini) and Ioannis Antonoglou (AlphaGo co-creator), launched March 2024, $2B raised. Then someone builds the comparison table against DeepSeek V4.1 Flash — 552B versus 501B total, 8B versus 23B active on prefill, 45T versus 28T pretraining tokens, and weights available on launch day versus later this month. Talk is cheap and this is a rather pointless announcement without anything backing it up, one commenter writes. He is right and it will still be the most-discussed model release of the week.
Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates — 475 points
475 points · 322 comments · vals.ai · HN discussion
Vals.ai published two candidate antiferromagnetic semiconductors for spintronic memory, found by a team of Claude Opus 5.5 agents — one compound designed from scratch, one that was first synthesized in 1999. Both are predicted to have zero net magnetism while still sorting electrons by spin, which is the property ordinary antiferromagnets lack and the reason MRAM researchers want them. The post shares the full calculations, the code, and a list of known caveats. That last item is the most credible thing in it.
The primer is where the trouble starts. A physicist commenter calls the introduction bizarre for presenting two kinds of magnet when everyone also encounters diamagnets and paramagnets, and another gives the diagnosis: this is what happens when Claude writes it and nobody reviews it. The substantive objection is from someone who did the work — theorists produced candidate high- and low-temperature superconductors constantly during his PhD, and fabricating the material is the hard part, layer by layer via CVD in a temperature-controlled room with a thirty-ton concrete dampener under the machine. Prediction is cheap; a predicted compound that nobody has made is a hypothesis with a DOI.
So the thread splits properly. One camp says after LK-99 this needs a truckload of salt, and points out the difference between having an idea and demonstrating the idea has merit. Another camp — the same one that calls LK-99 the most fun the internet has had in years — argues skepticism about “an LLM discovered X” is warranted while the underlying method is still worth tracking. The satire is the sharpest thing in it: a commenter reports that his own agent cluster discovered cold fusion last night, and notes he is barely at 13 percent of his weekly limit. Which is the actual finding here. The cost of generating magnet candidates has gone to approximately zero, the cost of validating one has not moved at all, and this post is a demonstration of the first fact standing in for the second.
Nobel Prize in Physics 2026: Francis Halzen — 444 points
444 points · 144 comments · nobelprize.org · HN discussion
The 2026 physics prize goes to one person, with a 1/1 prize share: Francis Halzen, University of Wisconsin–Madison, “for decisive contributions to the IceCube Neutrino Observatory and the discovery of high-energy neutrinos of astrophysical origin.” It is an instrumentation prize and it took 38 years to pay out. Halzen and J. Learned published the idea of using deep polar ice as both target and Cherenkov radiator in June 1988. The first sensors went into the glacier in 1992. AMANDA worked as designed but was too small. IceCube reached its full cubic kilometre in 2011 — 5,160 light sensors on 86 cables, at roughly two kilometres’ depth — with Halzen as principal investigator from the original 1999 NSF proposal that was funded in 2002.
The discovery itself was a lucky byproduct. In 2013, two petaelectronvolt events (1.04 and 1.14 PeV) turned up serendipitously in “starting event” data collected from May 2010 to May 2012, while the collaboration was looking for something else entirely at the exaelectronvolt scale. The most recent published result, from March 2026, includes a downgoing muon neutrino of 11.4 PeV recorded in 2019 — the highest-energy neutrino IceCube has published. IceCube is an NSF-managed facility with about 450 people from 58 institutions in 14 countries, and the prize citation is careful to name the collaboration’s work rather than implying a lone theorist.
The thread is better than the average prize thread because the corrections are the content. A commenter explaining Cherenkov radiation says particles travel faster than light; garyrob fixes it properly — faster than light in that medium — which is the difference between a physics error and a definition. Another asks how anything heavy reaches a construction site at the end of the longest logistics chain on Earth and gets the South Pole Traverse. A commenter who went down in 2009 to help with construction notes he did not see any neutrinos the entire time. And one describes a colleague who flew to the South Pole to install Debian on the data-processing machines, which is either the best or the worst job in science depending on your relationship with apt.
Polars 2.0 — 355 points
355 points · 80 comments · pola.rs · HN discussion
Polars 2.0 ships out-of-core spill-to-disk, a new Map dtype, stricter dtype handling and explicitness, a pile of optimizer work — join reordering, common-subplan elimination, dynamic predicates and bloom filters — and the strategic move: SQL as a first-class citizen. The post claims Polars now leads DuckDB 1.5.6, DuckDB 2.0 alpha and DataFusion 54 on TPC-H and TPC-DS at scale factors 10 and 100, run five times each in a hot setting, best-of-five, 60-second timeout, file cache cleared between engines, on a 16-vCPU machine and a 192-vCPU machine. The methodology is disclosed in more detail than most vendor posts, which is the most you can ask for in a benchmark where the vendor picked the queries.
The correct reading of it comes from a commenter who has worked on TPC benchmarking: never interpret a post like this as “database A is X percent faster than database B.” Read it as “we put focused work into the engine and expect certain workloads to improve.” That is not a knock. It is the only honest use of a vendor benchmark, and it is more useful than the charts. The thread’s practical content is better than the marketing anyway: one commenter reports the 1 Billion Row Challenge at 4m28s for pandas against 5.04s for Polars and 5.19s for DuckDB with DuckDB using 19 times less memory, another says Polars is a full pandas replacement for 99.9 percent of cases with geospatial as the exception (Geopolars is in development), and a third contradicts the benchmark entirely — DataFusion has always led Polars on their workloads with DuckDB far behind.
The unresolved question is the one that matters for adoption: SQL coverage is now first-class, but asof_join() is not reachable from SQL, and the class of “Polars-only” operations is exactly where the migration cost lands. A production user running billions of weather scores on the release candidate says it has been a lifesaver; that is one data point from someone who was already committed. The gap between “the engine is faster” and “your codebase moves” is not measured in TPC-H.
Example.com just launched the biggest redesign in decades — 318 points
318 points · 213 comments · debugbear.com · HN discussion
On 28 September, example.com — a domain IANA reserves for documentation and which has been a static English page for two decades — stopped being static. The new version cycles through explanations in English, Arabic, Chinese, French, Russian and Spanish every five seconds, using JavaScript: each character wrapped in its own span with an incremented CSS transition-delay, so the opacity fades character by character. By 3 October the animation was gone and all six languages render at once, after feedback that a five-second rotation is hostile to anyone reading slowly or using a screen reader.
IANA’s explanation, quoted in full, is about bandwidth: the site is heavily trafficked, so page content was split into a basic page augmented by a separate JavaScript file. Which is the funniest possible outcome — a placeholder whose entire purpose is to tell you not to rely on it, optimized for the traffic it receives from all the things that rely on it.
That is Hyrum’s Law stated as a case study, and the thread is the receipts. One commenter notes the immediate casualty: automated tests asserting on the page’s text. Another built a replica of the classic design on his own public endpoint-testing server so people can point their tests at something stable, open-sourced and self-hostable. A third explains the non-obvious reason example.com gets used in the first place — it does not present SSL/HSTS, which makes it one of the few clean ways to trigger a captive-portal detection path, and certificate pinning has broken the alternatives. My favorite is the person whose use case was a nearly white page for spotting smudges on his glasses, now ruined because the redesign renders dark by default in his browser; someone immediately offers a replacement. IANA’s “statement” in the article is the earlier email, edited, which one commenter finds disingenuous. Adding a language list to a placeholder broke a captive-portal tool, a glasses-cleaning routine and an unknown number of CI pipelines. Nothing else on this page is a better argument for treating every public URL as a supported API, or for changing a placeholder’s design on purpose once a decade so people learn the hard way.
The lamps in my house — 309 points
309 points · 117 comments · arslan.io · HN discussion
A tour of a designer lamp collection, with photographs, and no AI anywhere in it. The Artemide Tolomeo Mini (Michele De Lucchi and Giancarlo Fassina, 1987, Compasso d’Oro 1989), probably the most copied desk lamp ever made. The Tizio Micro (Richard Sapper, 1972) — no wires along the arms because the arms carry the current, counterweights so a finger moves it and it stays, and the small red details Sapper used as a signature on black products, the same instinct that later put a red TrackPoint on the ThinkPad. An Akari 1A and 3A, Isamu Noguchi’s washi-paper sculptures from a series he started in 1951 after visiting Gifu, where “akari” means both illumination and lightness. The author states plainly that none of them were cheap and that they are probably not worth the price on utility, that at some point this became collecting, and that vintage stores and fairs are where the good prices are. He also designed an adapter so his Tolomeo mounts directly to a USM Haller desk, which is the most engineer thing in the post.
The thread converts taste into a category. The best comment is from someone who points out that recessed ceiling lights do not establish hierarchy in the dimensions humans care about and that light at eye level is an accessibility and comfort issue — a former lighting designer confirms the term is task lighting and that the difference is not subtle. Then it does what HN does: one commenter says the $1,255 Muuto Post looks like dorm-room furniture, another notes the glass blocks in it are available for a few dollars and describes building an equivalent for under $15 with printed enclosures, and a third asks for a compact bright work light for fixing a 3D printer hot-end and gets three concrete hardware answers. A lamp post turned into a lighting-design seminar and then a buyer’s guide. This is what the front page is for, and it is the only story today where nothing is being announced, benchmarked or disputed.
Competitive Programmer’s Handbook (2018) [pdf] — 291 points
291 points · 71 comments · cses.fi · HN discussion
Antti Laaksonen’s handbook, 296 pages, draft dated July 2018, resurfaces on the front page for at least the fourth time. It covers the ground its title suggests — time complexity, sorting and binary search, data structures, dynamic programming, graph algorithms — in C++ with a companion problem set at cses.fi, and there is a commercially published, more polished Springer edition if you want the typeset one. The academic who comments that Laaksonen is one of his lecturers describes the courses as going into the bread and bones of CS and leaving you feeling 10x smarter, which is a better endorsement than the download count.
The reason it is worth more than a line today is the top comment, which has nothing to do with the book: a developer says he has been doing Codeforces, Rosalind, Codewars and LeetCode in his spare time as a detox from heavy daily agent use, relearning things that over-reliance on LLMs made him forget, and that he missed the feeling of thinking hard and arriving at a solution. He hopes we do not lose that desire as a species. The reply is the counter-argument in one line: the rise of LLMs has demonstrated that most people never had the desire, and only a minority ever did.
That exchange is the most honest artifact on today’s page about what all of this is doing to skill. Not “will AI replace programmers” — nobody here is arguing that — but whether the specific pleasure of solving a hard problem survives when a tool is faster at it. The second-order observation, which nobody in the thread makes explicitly: the handbook’s own genre exists because practice matters, and the practice is the part being outsourced. Notably, the commenters who argue the book is too terse for learners recommend alternatives — Zingaro’s Algorithmic Thinking — which is a scarcity-of-attention problem, not a scarcity-of-material one.
Friendship ended with Deno, now Node is my best friend — 291 points
291 points · 213 comments · dbushell.com · HN discussion
David Bushell went back to Node for a SvelteKit client project after years on Deno, and the surprise is that the things he left over no longer exist. require() is gone from his code, the modern APIs are supported, fnm handles version switching, and Node runs TypeScript natively — with one philosophical exception, quoted from the Node documentation: type-stripping is refused for files under node_modules, because the project wants to discourage authors from publishing packages written in TypeScript. The error code is ERR_UNSUPPORTED_NODE_MODULES_TYPE_STRIPPING. His reading is that this is Microsoft-flavored ecosystem pollution control, and he is clear that it is a policy, not a limitation of the runtime.
The supply-chain paragraph is the part worth stealing. He switched to pnpm to avoid getting immediately pwned, set minimumReleaseAge: 1440 and trustPolicy: no-downgrade, notes pnpm blocks post-install scripts while NPM still runs them, and says he settled on a one-day release delay because a month was long enough to break dependency resolution. “Long enough to allow some other sucker to beta test the next Shai-Hulud” is the correct framing for the entire npm worm era: the defense is not review, it is a delay and hoping someone else is the canary. Every dotfile he needs is one more mistake, he writes, and he is not thrilled about the two tsdown added.
The thread’s best information is about Deno’s decline, from someone who contracted on the standard library: after the layoffs there is no roadmap and no communication, and it is sad to watch. A reply argues Deno’s only sensible play was to get acquired, and that you either sell or get squeezed out. The pushback is that Deno still wins on the boring parts — built-in test runner, linter, formatter, type checker, JSR for publishing — and that Node has no obvious answer for any of them. That is the shape of the runtime war in 2026: Deno’s ideas won, and Node got them by absorbing rather than losing, which is how the browser wars always ended too. The commenter who says the post’s negativity cost it credibility is right, and it does not change the conclusion.
Find the flattest route between any two points in SF — 287 points
287 points · 100 comments · flattensf.com · HN discussion
Drew Edwards built a route finder whose objective is not distance but cumulative climbing. It runs in the browser over 160,000 street segments using USGS one-metre lidar elevation and the Overture/OpenStreetMap graph, and instead of returning one route it exposes a slider along the Pareto frontier between shortest and flattest — every route nothing else beats on both counts, from the fastest path to the one worth walking, where a foot of climb is priced at 200 feet of walking. Sliding right never shortens the route and never adds climbing. On foot, stairways are allowed; on a bike they are excluded. There is a “prefer calm streets” mode that measures distance in comfort using the SFMTA bikeway network: a protected lane counts as 0.8 of its length, a quiet street as 1, a busy arterial without a lane as 1.4 to 2. Place search runs offline against data baked into the page.
Two things make this the best-built toy on the page. The first is that the objective function was chosen deliberately and the trade-off is exposed rather than hidden — “flattest” is under-specified, and rather than guessing, the author gives you the whole frontier. The best objection in the thread is exactly that: a commenter wants grade minimized rather than total gain, so that SOMA to Nob Hill approaches from the east or west instead of attacking Taylor Street head-on. The reply is the correct counter — minimize maximum grade and you get infinite zig-zagging up a mountain, because no finite path is ever flat enough. The second is the maintenance loop: a commenter reports a wrong route from Cabrillo Street, the author diagnoses it within the thread, finds that San Francisco’s Slow Streets “destination-only” tag was being parsed as closed to pedestrians, fixes it and deploys, all in public. That is the good version of shipping software in front of an audience.
The rest of the thread is the usual mixture of route quibbles and city politics. A commenter is told to bike on Geary and Divisadero and declines on the grounds that he would rather stay alive; the author concedes the tool does not know bike lanes yet and says he will pull the SFMTA network. Someone asks for the same thing in Istanbul, which is a genuinely hillier problem. Somebody points out that 25-centimetre 3DEP elevation data is available free on S3 and would fix the DEM accuracy complaints, which the author does not dispute. The tool is not perfect at what it claims. It is honest about the trade-off it is making, which is why people are arguing about the second decimal place instead of dismissing it.
Nature’s capacity to ‘bounce back’ when species are lost is overestimated — 264 points
264 points · 129 comments · phys.org · HN discussion
The largest synthesis of its kind, published in Nature Ecology & Evolution and led by King’s College London with Imperial College London, the Natural History Museum and the Alan Turing Institute: 423 studies, 222,829 data points, 23 categories of ecosystem services and functions, land, freshwater, marine and estuarine — a database more than twice the size of the previous best effort. The finding is a negative one. Across 23 categories, most benefits keep rising with biodiversity rather than levelling off once a handful of species are present, which means functional redundancy — the assumption that species can substitute for one another and losses do little harm — has been overestimated. Ecological benefits climb as diversity rises, so they fall as it drops.
The paper’s own caveats are the reason to take it seriously. Oceanic carbon sequestration shows by far the strongest sensitivity of any service measured (Fisher’s z of 1.48 against −0.03 for air-quality regulation), and that result rests on seven datasets against 154 for terrestrial carbon sequestration — which the authors flag as an urgent gap rather than burying. Marine life cannot be an afterthought if the world is counting on blue carbon, the lead author says, and the honest version of that sentence is that we have seven studies on it. The counterexample is equally instructive: hazard regulation appears insensitive to biodiversity (r² of 0.17) because dune stabilization is done by one or two foundational shrub species. Diversity is not uniformly protective; some services are hostage to specific taxa, and the paper says so.
The thread’s strongest contribution is skepticism about the equilibrium metaphor itself. A commenter recommends Adam Curtis’s All Watched Over by Machines of Loving Grace and argues that the idea of a natural equilibrium is a 1950s cybernetic fantasy projected onto ecology: model the world in systems language and you can claim the system will return to equilibrium if you just wait longer. Another commenter warns against the mirror-image error — the “models are wrong” reflex curates counterexamples and blinds you — which is the right caveat to attach. Then the cod fishery arrives as the empirical case: after the collapse, regulation let the population come back, it did not, and the niche filled with other species. As one reply puts it, that is a story about nature bouncing back, just not the way the people managing it thought. Nothing here disputes the paper. It disputes the comfortable word “back.”
Dust: Pretraining Transformers Without Backpropagation — 257 points
257 points · 73 comments · qlabs.sh · HN discussion
Dust is a zeroth-order training method that perturbs activations — node perturbation — independently at every token, so each token acts as a virtual population member and one forward pass evaluates them all in parallel. The claimed results: gradient estimates that align with backprop and improve as the population grows, sometimes exceeding it, up to one billion tokens of testing; larger models being more population-efficient rather than less, with a 243M-parameter model beating one 120 times smaller at most population sizes; and 10³ to 10⁴ times better efficiency than weight-space evolutionary strategies above a million tokens, based on extrapolation. The framing is Bitter Lesson maximalism — differentiability and backprop as low-compute inductive biases that stop being worth their constraints when compute is abundant.
Read the fine print, because the fine print is the argument. Dust approximates backprop “at large population (i.e. substantially more compute)” and is orders of magnitude more efficient than one specific baseline, EGGROLL, by extrapolation. Those are not the same claim as “competitive with backprop,” and the post is more scrupulous about this than its title. The extrapolation is doing real work: the efficiency multiples are projected, not measured, at the scales where any of this would matter.
The thread is the right kind of hostile. A commenter who has watched derivative-free optimization get hyped every few years says he would bet his life savings that none of them ever makes an impact, because neural network objectives are smooth and the gradient is a strictly better signal than random directions. The sharpest reply is that this is not an either/or: gradients can be computed numerically, so any method that samples the cost function can recover a gradient when it needs one. A third objects that the Bitter Lesson framing does not hold up, because zeroth-order methods do not address nonconvexity — Dust smooths the landscape, and smooth rules already apply to first-order methods, so remove nonconvexity and the advantage evaporates. The best steel-man in the thread is from someone who sees zeroth-order work not as a replacement but as a technique for regimes where backprop is structurally weak, and who proposes using activation-space search to generate teaching targets for backprop. That is the version of this research that will still matter in five years, whether or not Dust scales.
Gleam doesn’t compile to Erlang source anymore — 252 points
252 points · 103 comments · gleam.run · HN discussion
Gleam v1.19.0 replaces its Erlang code generator with one that emits Erlang abstract forms instead of Erlang source — the metadata-annotated IR the Erlang compiler’s parser normally produces, binary-encoded via the external term format, which lets Gleam load generated code into the compiler’s back half directly. The gains are concrete: faster compilation, and stack traces and BEAM crash reports that point at the original Gleam line rather than the nearest function boundary, because the location metadata now describes Gleam source. It also puts full debugger support within reach of a third party, and it ends Louis Pilfold’s stated ambition that nobody ever gets to use “transpiler” as a pejorative again.
The benchmark is 100 modules of 100 functions, from José Valim’s langcompilebench, and the post says outright that benchmarks are contrived and this shape only exercises a subset of the language. That caveat is the reason to trust the direction and not the number.
The single most useful comment is a correction: the title makes it sound like Gleam stopped producing something Erlang can use. It did not — the output moved to a lower-level format, and the same IR is what Elixir compiles to and what Erlang’s parse transforms manipulate, so anyone who has written one already knows how to read Gleam’s output. The author confirms in-thread that LFE has moved the same way. The rest of the thread is a community developing the case for the language: someone maintaining 200,000 lines of Gleam says the Discord has tolerated even his stupidest questions; a commenter praises the contributor who did the rewrite for streaming the work publicly without ego; and the requests are all about reach — a native target, wasm and WASI. For a language whose pitch is that it is small and boring on purpose, this release is the unglamorous kind of progress: a compiler change nobody outside the BEAM ecosystem will notice, that makes every future stack trace correct.
Best-selling Zigbee temperature sensors tested and compared — 206 points
206 points · 112 comments · smarthomescene.com · HN discussion
Twelve best-selling Zigbee temperature and humidity sensors — Aqara, Sonoff, Tuya, ThirdReality, Zemismart, Moes, sourced mostly from AliExpress — tested head to head against a single calibrated reference. The reference is a Xiaomi Miaomiaoce MHO-C401 running pvvx custom firmware, calibrated against a lab-certified thermometer that lives in a butcher’s commercial refrigerator and is checked by the state every three months. That required a −0.5°C and +4 percent humidity offset to match, and the author is explicit that this anchor was taken at roughly 4°C, so it does not guarantee straight-line accuracy at room temperature — what it guarantees is a consistent reference for comparing sensors under identical conditions, and the sensors can be recalibrated in Home Assistant anyway. Testing was done on a fresh Zigbee network with a PoE coordinator, 3.5 metres of separation, no other paired devices, Zigbee2MQTT 2.12.1-dev.
This is the highest-value hardware methodology on the page precisely because it does not claim more than it can support. Twelve units from six brands, no displays, all battery-powered; the differentiators are accuracy, range, build, battery type and form factor, and the follow-up added refrigerator benchmarks at the temperatures where cheap sensors actually fail.
The thread then argues about what the review should have measured. One commenter wants the sensor die and the MCU disclosed — Bosch, Sensirion, thermistor — because those determine accuracy and battery life. The reply is correct in the abstract: buyers cannot verify what is inside a unit anyway, and unit-to-unit variation inside one SKU is larger than the difference between Chinese and Western parts. The more practically useful notes are about batteries and lifecycle: IKEA’s sensors take AAA cells instead of coin cells, one commenter’s IKEA Thread units have mediocre battery life and no configurable update rate, and everyone who has put a cheap sensor inside a fridge agrees it is the killer app. One commenter tracked a fridge that had been quietly sitting between 10 and 15 degrees for half a summer; another tracked down a box that was preventing a door from closing. A $5 sensor and a notification rule caught failures that a decade of opening the door had not. That is the case for this whole category, and it does not need a benchmark to make it.
Also on the page
Kolibri: A Sovereign Open-Weight Model — 699 points · 336 comments · aleph-alpha.com — the one announcement today with weights attached: a 78B-parameter, 3B-active English-German mixture of experts with a 1M-token context, downloadable now from Hugging Face under Apache 2.0, built on a pipeline validated by an earlier 30B model and hundreds of ablation runs with unattended pretraining that survives hardware failures. The interesting part in the thread is that the technical report reads like a tutorial on building a modern agentic LLM, including how the dataset was constructed. The two corrections worth keeping: a commenter points out that a post this invested in sovereignty should mention the company’s slated merger with Cohere, since these non-US, non-Chinese labs increasingly have to share effort and cost; and a benchmark claim from Kolibri’s own harness, on Kolibri’s own benchmark, has Qwen 3.8 27B beating it 79.9 to 70.8 in German. Kolibri is trained with abstention data so it says “I don’t know” when the answer is not in context, which is a more useful product feature than most of the numbers above it.
JetBrains reports revenue growth, net financial loss for 2025 — 553 points · 505 comments · helgilibrary.com — CZK 16.0 billion of revenue, up 6.3 percent and a record, against a net loss of CZK 315 million, a −11.1 percent return on equity and a 5.73 percent EBITDA margin, with CZK 7.6 billion of net cash on the balance sheet. The thread skips the loss almost immediately and argues about the strategy instead, which is right: a company with net cash and a record top line investing into a loss is not a distress story. The accusation that lands is that JetBrains saw the agentic IDE coming and staffed it with two people and some internships, and the counterpoint is from people who still use Claude inside a JetBrains terminal because diffs, navigation and worktrees matter. What nobody in 505 comments can point to is the IDE that does deep worktree support, parallel sessions and diff review better than a command line, which is either a large open product surface or a verdict on the category.
Pi Durable — 510 points · 74 comments · earendil.com — Earendil shipped Pi 1.0 on 1 October and immediately followed it with an experimental harness for agents that must run for a long time and survive failure: reachable from different surfaces, infinitely long conversations, recovery from catastrophic internal and external failure, multiple humans steering the same agent. It is explicitly not a replacement for the Pi coding agent but a framework sharing its pi-ai code and its stated principles of minimalism and malleability. The thread names the competitive set — LangChain Deep Agents, Vercel Eve, OpenAI’s Agents API, Anthropic’s managed agents — and notes the design choice that the release notes do not advertise: Durable drops branching conversation trees in favor of forks carrying ancestry information. Whether that is required for the durability guarantees or a step back from the coding agent is the open question, and it is the kind of question that decides adoption six months from now.
ChatGPT is adding real cartoonists’ signatures to fake New Yorker cartoons — 533 points · 388 comments · niemanlab.org — a viral cartoon about Dolly Parton and Tim Curry was generated by ChatGPT from the prompt “a New Yorker-style cartoon” and shipped with “BLOPER,” the pen name of real cartoonist Brendan Loper, in the corner. Loper did not draw it or sign it; a tweet of it got 25,000 likes; Nieman Lab commissioned him to draw a cartoon about the experience. The technical explanation in the thread is the right one — training data associates that style with that signature in the corner, and nothing makes a model treat a signature as special unless someone trains it to — and the legal intuition is the correct framing: a human forging another artist’s signature would be liable, and the same should apply to whoever generated it. Gwern notes he erases the false signature with an extra edit and that most users will not bother, which is the mechanism by which this scales. An engineering manager reportedly put one in a sprint demo.
AI is now capable of developing its own inference hardware — 159 points · 176 comments · github.com/FeSens/openTPU — openTPU is an open-source AI accelerator in one readable monorepo: SystemVerilog hardware, an ISA, a bit-exact simulator, a kernel language and compiler, and host software for a real PCIe card. It runs ten models with their real weights on an Inspur YPCB-00338 card (a Xilinx Kintex-7 xc7k480t with two DDR3 channels) and produces tokens bit-identical to the simulator. The headline is doing more work than the artifact — a commenter’s deflation is precise: a human prompted an LLM to build a simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize designs against constraints in that simulation. The best technical objection is also in the thread: someone opened the RTL, saw the floating-point math, concluded that correct floating-point operations apparently are not required for LLMs, and closed the page. A Kintex-7 and DDR3 are a decade-old board. The achievement is the loop and the readable toolchain, and it should be sold as that.
World’s first enhanced geothermal power plant completed in just 23 months — 134 points · 54 comments · techcrunch.com — Fervo started selling power from Cape Station in Utah on 30 September, a day early, with the first third of a 100MW plant synchronized to the grid and a site potential the company has put at 4GW. Groundbreaking to commercial operation took 23 months and the target for future blocks is 18, which matters because phased development matches how data centers are built. Google and Southern California Edison have committed to buying power. The thread removes the marketing: TechCrunch reads like a summary of a press release, there are dozens of enhanced geothermal projects, and two Swiss attempts in Basel and St. Gallen were abandoned after causing earthquakes. The number that actually decides whether this replicates is thermal drawdown — if the fractured rock cools faster than modelled, the operator has to keep drilling — and this plant has no operating history yet.
Tapo (Rust/Python library) now speaks TP-Link’s TPAP protocol — 107 points · 45 comments · dinculescu.dev — since firmware 1.4.0, TP-Link’s Tapo plugs and bulbs have refused third-party clients unless the user flips a “Third-Party Compatibility” switch buried in the app, which is the third time in three years a security-motivated firmware change has locked out unofficial clients. As of v0.11.1, the unofficial Rust/Python client works with that switch off. The historical context is in the post: a 2023 Catania/Royal Holloway disclosure on the L530E bulb protocol triggered firmware changes that broke integrations, and users experience all of it as “my script returned 403 and nothing in my code changed.” The commentary is more interesting than the library. One commenter argues reverse-engineering closed protocols became cheap around GLM 5.2, because a binary and Ghidra’s MCP integration is now enough to get most of the way — and the same model capability is why another commenter calls the writing style detectable AI slop and refuses to read it. Both of those are load-bearing facts about the ecosystem in 2026.
Erdosproblems.com Succumbs to the AI Onslaught — 29 points · 3 comments · erdosproblems.com — tiny story, disproportionate signal. Thomas Bloom built erdosproblems.com to promote a list of problems to human mathematicians; it became an accidental AI benchmark because the problems are easy to state and range from obscure to deep. He now describes the effects: some mathematicians dismiss the field as recreational, others have stopped working on Erdős problems because they believe they cannot compete with AI, and — the part that pushed him to write — the main way people interact with the site is now to advertise AI-generated proofs with no explanation, as a way to stake an increasingly meaningless priority claim, which is not the site he wanted to run. He is asking for feedback on what to change. The HN submitter’s clickbait title is the least interesting thing about it.
Benchmark in Milliseconds — 98 points · 22 comments · matklad.github.io — Alex Kladov’s rule of thumb: tune a microbenchmark’s input size until it takes about 300 milliseconds, because three-digit integers are precise enough to see improvement and easy to scan, hundreds of milliseconds makes fixed overheads irrelevant, the number is human-perceptible so you can use intuition instead of numeracy, and anything over a second makes iterating on the benchmark too slow. His stated assumption is that the purpose of benchmarking is not precise measurement but enough intuition to make a correct decision. The thread’s pushback is the standing critique of the whole genre: report confidence intervals, compare against a control in the same run and round-robin, or use a framework that handles warmup and stabilization — because on real hardware, clock scaling, interrupts and thermal throttling produce irreproducible numbers if you do not. Both positions can be true: the right answer depends on whether the number you want is a decision or a claim, and today’s front page is full of people making claims.
Still on the page
Ninety of the 124 stories above 200 points were covered in earlier roundups. Deltas are against the last roundup that tracked each story, which for most is yesterday’s.
Three of them moved enough to matter, and two are the ones nobody is looking at. Anthropic reported diary entry to police 315 → 795, up 480, is by a factor of four the biggest mover in the window, with 630 comments, and it sits at rank 36. Web Search API 391 → 581, up 190, is the second, at rank 41. Denmark’s 8.8M-person data breach 428 → 496, up 68, at rank 63. Then a middle cluster: Qwen 3.8 Flash Next on a 4090 904 → 921, up 17, after yesterday’s 473-point jump; Google Data Center water 504 → 526, up 22; Pixel 11 and GrapheneOS 363 → 426, up 63; Apple Intelligence removal 738 → 765, up 27; Budget caps 623 → 636, up 13; Bob Cringely 926 → 936, up 10; the electrician essay 469 → 478, up 9; and Kolibri 692 → 699, up 7.
Everything else moved by six points or fewer, across nearly eighty stories. The top of the list barely breathed: Gemini 4 Argon 1,699 → 1,701, Pi 1.0 1,684 → 1,686, Utah’s VPN ruling 804 → 808, You said no MCP 682 → 686, Mike Tomlin’s Minecraft city 672 → 675, Clef 638 → 640, StreetComplete on iOS 629 → 630, Git 3.0’s SHA-256 default 573 → 576, Apple Pass Designer 566 → 567, Frog and Toad 584 → 587, the Linux kernel advisory 576 → 579.
The exception that proves the rule is the shape of the two big movers: both gained hundreds of points at ranks 36 and 41, which means they accumulated them somewhere the front page does not display. The story with 795 points and 630 comments — the most-discussed thing in the window apart from Mistral — visited the front page and left. Meanwhile the page today is topped by a model announcement with no weights, a 1987 desk lamp and a placeholder domain’s CSS transitions. Rank is a decay function and points are a cumulative one, and the gap between them is now large enough that a roundup which reads the front page and a roundup which reads the rank list are describing different days. This one is reading the rank list.
Throughline
First: the open-weight frontier has become a genre of announcement rather than a distribution channel, and today both of its leaders shipped the genre. Mistral’s trillion-parameter model with 49 billion active parameters arrives as a public preview with the weights at month’s end; Reflection’s 501-billion-parameter Beam arrives with 100 million RL rollouts on 10,500 GPUs and the weights “later this month.” The only completed open-weight release on the page is Aleph Alpha’s 78B Kolibri, which nobody argues is frontier capability and which is the one being used in production, in regulated industries, in German. The comments write the verdict themselves — publish your weights or shut up — and the fact that this is now the standard shape means the loading screen, not the artifact, is what gets benchmarked, compared against DeepSeek tables and argued about for 164 comments. Mistral’s pre-release plan is the one genuinely new thing in either announcement: it will hand an unlocked, cyber-specialized trillion-parameter model to outside firms and state authorities before the public gets the locked one. That is either the most responsible way to ship dangerous capability or a permanent structure for distributing it, and nothing in either document says which.
Second: today’s page is a referendum on validation, and every single vote went the same way. Dust claims backprop-free pretraining competitive at large population — meaning substantially more compute — and 10³ to 10⁴ times better efficiency than one specific baseline, by extrapolation; the thread’s response is that gradients are a strictly better signal and the numbers scale by projection. Vals.ai’s Opus 5.5 agents designed two antiferromagnetic candidates and published the calculations, and the response from someone who has actually done the work is that theorists produce candidate materials constantly and fabrication is the hard part. Beam’s headline generalization demo is audited within hours and survives only in weakened form. Polars leads DuckDB and DataFusion on a benchmark the vendor wrote, and a TPC veteran explains that the correct reading is “we worked on the engine,” not “we are faster.” Gleam benches a rewrite on 100 modules and says up front that the benchmark is contrived. The generative cost of a candidate, a proof, a benchmark table or a model release has gone to approximately zero, and the cost of one validated result has not moved at all. What has changed is that more of the page than ever is the theory, and the theory is now indistinguishable in form from the result unless you read it carefully. The one item on today’s page that resolves this honestly is Erdosproblems.com — a site that became a benchmark by accident, whose maintainer is now dealing with a flood of unexplained AI-generated solutions filed as priority claims, and asking what to do about it. That is the problem the other 123 stories have not solved.
Third: the smallest engineering decisions on this page are the ones with consequences, which is the most consistent thing about 2026. A placeholder domain split its content into a basic page plus a JavaScript file to save bandwidth, and in doing so broke CI pipelines, a captive-portal trigger, and a man’s method for spotting smudges on his glasses. A language compiler moves from emitting Erlang source to emitting Erlang’s internal representation, and every future stack trace in the ecosystem points at the right line. A 4°C calibration offset from a butcher’s state-inspected refrigerator produces the most trustworthy hardware review on the page. A firmware switch named “Third-Party Compatibility” turns every unofficial Tapo integration into a 403 with nothing in the logs. A fridge sensor and a notification rule catch a cooling failure the door had been hiding for half a summer. None of these were announced with a benchmark or a parameter count, and all of them are load-bearing. The throughline of the day is not that AI is arriving or failing; it is that the front page’s supply of announcements is now larger than its supply of things that work, and the things that work are quietly one level down, in the middle of a thread about lamps.
Published 6 October 2026. Scores captured at 13:05 PDT from the HN Firebase API — the front-page set at 13:04, the full 500-item top-stories rank list at 13:07 — with deltas taken against the score quoted in the most recent roundup that covered each story. No Algolia sweep today; off-window scores are from the rank list. The front page keeps moving after that.