HN Brief: 2026-08-11

Today's HN was dominated by the friction between open and closed systems. Meta's new local-agent model sparked enthusiasm, but the celebration curdled into arguments about server scraping and data surveillance. A security researcher reporting a 180,000-meeting Firebase leak was ignored for six months, sparking a debate about whether naming exposed clients was warranted or reckless. Meanwhile, Illinois passed a law that could technically put Linux distros on the hook for age verification, and the UK's push to export anti-anonymity laws to the US split commenters between "think of the children" skepticism and genuine worry about algorithmic harm.

Jump into “Tl;dv: Over 180k meetings left wide open” for the infuriating story of a researcher who joined live government calls while the CEO sat on the fix. “Illinois just passed a law that puts Linux on the hook for age verification” is worth clicking for the systemd flamewar that revealed hobbyist distros might be technically liable. “Learning more about Claude's mathematical capabilities” has the bizarre tale of an AI being told “believe in yourself” for 650 failed ideas before cracking a number theory proof. “Exploiting System Management Mode with a very long interrupt” is a delightfully absurd hardware exploit where one slow MMIO read cracks open CPU firmware. And “Humanising LLM Outputs Is Dumb” captures the collective vent about Claude Opus 5's flowery, unreadable prose.

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows [comments]

1103 points · 603 comments · research.meta.ai · 21h ago

Meta released Muse Glimmer, a 30B-parameter open-weights model aimed at running agents entirely locally on a single consumer GPU. Hacker News was pleased to see Meta staying open-weight, but several people immediately pushed back with complaints about Meta scraping their servers for training data despite cease-and-desist requests. The thread quickly turned into a head-to-head comparison with the upcoming Qwen 3.6 27B and Google’s Gemma 4, with split opinions: Qwen overthinks in multi-turn tool calls (especially with reasoning on) while Glimmer’s quantized 17GB variant reportedly only loses 1% on benchmarks. A side conversation about using Qwen as a TTRPG game master revealed a real hunger for these local agents, but also pragmatic gripes about memory constraints and the need for tool-use harnesses that don’t bloat context.

Tl;dv: Over 180k meetings left wide open [comments]

575 points · 191 comments · bobdahacker.com · 19h ago

The linked article details a security researcher's discovery that tl;dv, an AI meeting recording platform, left its Firestore database wide open for over six months, allowing any authenticated user to query over 180,000 meeting records and even join live calls uninvited—he walked into a Malaysian Ministry of Education meeting and a US university startup call. The thread largely focuses on the sheer negligence of the company ignoring the researcher for six months, with many arguing that the CEO’s inaction after being explicitly told makes the breach far worse than a simple coding mistake, calling for a full shutdown. A strong split emerges around the researcher’s decision to name specific clients (government agencies from 23 countries, universities, and companies like HubSpot), with some saying it unnecessarily exposes them to risk while others counter that the risk already existed, the fix was ignored for half a year, and public shame was the only remaining lever. There's also a persistent undercurrent of criticism aimed at Firebase itself, with several people arguing the platform makes it too easy to ship insecure default configurations, and that this is just the latest in a long line of similar exposures. A handful of people selling competing AI note-takers show up to pitch "local" alternatives, but the main takeaway is that a company with SOC2, GDPR, and EU AI Act compliance badges let an attacker sit in on live government meetings for half a year because no one checked a single collection's security rules.

Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models [comments]

488 points · 446 comments · www.ft.com · 18h ago

Mark Zuckerberg published a manifesto arguing that Meta is returning to its roots by betting on open models and attacking "closed" AI rivals like OpenAI and Anthropic. The HN thread immediately read this as a classic strategic move—when you're losing the race, you try to change the rules by making your competitors' closed moats look anti-competitive. A lot of skepticism came from users who pointed out the contradiction between "open models" and Meta's business model of hoovering up personal data; one side argued that open-sourcing the model is a genuine gift that lets you self-host and avoid Zuck's data-collection entirely, while the other side shot back that the whole pitch is just a way to get developers onto Meta's compute infrastructure and normalize surveillance as a feature for the "personal agent." The thread also dug up the old "dumb fucks" quote about user trust, splitting into people who see it as a timeless warning about Zuck's character and people who insist it was just a sarcastic 19-year-old being flippant about how little anyone cared about privacy. A smaller tangent spun off into whether the manifesto itself was LLM-generated, based on its use of double hyphens instead of em dashes, but that got shot down quickly.

The UK's war on anonymity has come to America [comments]

461 points · 343 comments · www.effort.news · 8h ago

The article details an investigation into a coordinated push by British NGOs—including 5Rights, the Center for Countering Digital Hate, and the Institute for Strategic Dialogue—to export UK-style digital ID and age verification laws to the US, using child safety rhetoric to justify measures that effectively end internet anonymity. The thread immediately zeroes in on the core tension: commenters split sharply between those who see any "think of the children" argument as a transparent power grab that should be dismissed outright, and those arguing that tech people refusing to engage with genuine parental concerns about online harm is what guarantees even worse regulation down the line. The parental controls debate becomes a flashpoint—one side points out that technically savvy kids have always bypassed controls, while the other retorts that today's commercial internet, weaponized by trillion-dollar corporations and dangerous influencers like Andrew Tate, is fundamentally different from the wild-west bulletin boards of the 1990s. A deeper split emerges between people who want to regulate the actual harmful platform behaviors (addictive algorithms, dark patterns) instead of mandating ID, and those who say that's politically unsalable without a similarly simple, concrete alternative. The underlying worry that runs through the whole thread is that even well-intentioned filtering mechanisms will inevitably be turned on political dissent, as the article claims has already happened in the UK.

Mars Bar from 1991 found – and it's 20g bigger than today's [comments]

339 points · 498 comments · www.bbc.com · 16h ago

A house clearance in Scunthorpe turned up a Mars Bar from 1991 that’s 22.5 grams heavier than the current 40g version, which the BBC article uses as a prop for the old shrinkflation chestnut. The thread immediately split between people shrugging that smaller junk food portions are probably a net public health win and others pointing out you can just eat two, so the calorie reduction is mostly a fiction unless the price fell proportionally—which it didn’t. Several deep dives into the history of Mars mentioned that the UK bar was actually invented by Forrest Mars after he split from his father’s American company, and that the 2002 recipe was a deliberate shift to a lighter, Milky-Way-style product, not a sneaky downsizing. A few commenters had fun with the Scunthorpe location, noting the town’s famous name-filter problem, while others dismissed the whole thing as silly-season filler. The underlying split was between people who see shrinkflation as proof of corporate greed and those who think paying more for less sugar in a candy bar is a weird thing to be mad about.

Illinois just passed a law that puts Linux on the hook for age verification [comments]

315 points · 448 comments · linuxstans.com · 11h ago

Illinois passed HB5511, the Children's Social Media Safety Act, which on its surface targets algorithmic feeds and nighttime notifications for minors on platforms like TikTok and Instagram, but buried in the text is a separate set of requirements for "operating system providers" to build an age-declaration system and hand an age-bracket signal to any app that requests one by 2028. The HN thread immediately zeroed in on the glaring problem: unlike Colorado and California, which both amended their similar laws to carve out open-source software, Illinois included no such exemption, meaning a hobbyist Linux distro or a project like systemd could technically be on the hook. A huge chunk of the discussion got sidetracked into the usual systemd flamewar, with people pointing out that systemd already stores user birthdates and joking that this finally gives the haters a legitimate grievance, while others debated whether "algorithmic feed" is even coherently defined in the bill's carve-outs for chronological and manually curated content. The enforcement reality got a reality check: only the Illinois Attorney General can sue, there's no private right of action, and the penalty caps are $7,500 per child, not the $50,000 the governor's press release advertises, so most agree that only big players like Google, Microsoft, and Red Hat with Illinois revenue will actually feel pressure, while the rest can laugh it off.

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots [comments]

306 points · 111 comments · cactuscompute.com · 14h ago

The linked article is a detailed technical post from Cactus Compute introducing Needle2, a 14MB agentic LLM (45M parameters at 2-bit compression) designed exclusively for tool calling, structured extraction, and device control on ultra-low-power hardware like Raspberry Pis, microcontrollers, and sub-$200 phones. Hacker News was genuinely impressed by the engineering—hitting 500+ tokens/sec on a Pi 5 and fitting in 28MB of RAM is no joke—but the real test came from people actually trying the web demo, which exposed a critical weakness: the model hallucinates tool calls even on utterly out-of-domain input like the word "potato," confidently producing a `lock_door` invocation with a confidence score of zero. The team’s defense is that you're supposed to threshold on that learned confidence score and escalate to the cloud when it's low, but several commenters pushed back, correctly pointing out that calibration isn't guaranteed and that without proper out-of-distribution detection benchmarks, a 14MB model that confidently guesses "front door" from noise is a liability, not an asset. Others debated whether the right architecture is a single tiny model or a hierarchy of them, and a few folks dug into the novel "engram" hashed-n-gram tables and sliding-window attention, questioning how those hold up under real-world abuse versus the paper’s claims.

Squeak 6.1 [comments]

260 points · 126 comments · squeak.org · 19h ago

Squeak 6.1 "Vanessa" landed after four years and 1,700 patches, adding a new tree browser, a revived Objectland, and deep kernel rework for process scheduling and class reshaping. The thread quickly pivoted from the release to the broader philosophy of Smalltalk, with many arguing it rewires your understanding of what object-oriented programming actually means—and that nearly all of JavaScript’s good parts trace back to Smalltalk. That claim got immediate pushback: Self, not Smalltalk, gave JS its prototype system, and several people argued that prototypical inheritance has been a security headache, not a blessing. The conversation also lit up around the idea of a persistent, live image you sculpt over years—some called that way more practical than the tear-down-and-restart model of most languages, while others pointed out that Common Lisp, Emacs, and even APL already do the same thing, and then the thread split into a debate about whether NixOS is actually the opposite of a live image (immutable, compiled artifact) versus Debian’s mutable-pile-of-state approach, which feels much closer to Smalltalk’s live environment.

Stop Killing Games: It's time to sue Sony, join us [comments]

230 points · 126 comments · www.massaschadeconsument.nl · 11h ago

A Dutch consumer foundation is suing Sony, arguing that its closed PlayStation Store constitutes an illegal monopoly that lets it overcharge for digital games—the "Sony Tax"—compared to physical copies sold elsewhere. The Hacker News thread quickly split: some commenters dismissed the case as nonsense, pointing out that PlayStation isn't a monopoly because you can buy an Xbox or PC instead, and that the real problem isn't store exclusivity but the lack of ownership when Sony can revoke digital licenses. Others pushed back hard, arguing that the high upfront cost of the console plus a growing library of digital purchases creates lock-in, and that the comparison to Apple's mandatory App Store is apt—especially since EU regulators already forced Apple to open up. A side tangent emerged around "right to unlock" the hardware to install alternative OSes, though several people noted that even if you got Linux running, Sony would cripple GPU access, making it pointless.

Parametron: 50s Japanese computer that uses neither transistors nor vacuum tubes [comments]

221 points · 53 comments · ethw.org · 21h ago

The article celebrates the 1954 invention of the parametron by Eiichi Goto at the University of Tokyo—a logic element using cheap ferrite cores and parametric oscillation that powered Japan's early computers like the PC-1. The thread quickly pivots from the achievement to the hard technical trade-offs: parametrons ran at 10-15 kHz versus MHz for vacuum tube computers, and their power draw shot up with speed, making them physically bulky and a dead end once discrete transistors arrived. A major split emerged between people arguing the parametron was a clever, budget-constrained solution that enabled Japan's postwar computing industry and those insisting it was a fundamentally unscalable technology abandoned for good reason. Some commenters dive into the quantum flux parametron as a modern superconducting descendant that could reach GHz, noting it's a more promising alternative than current quantum computers, while others point out the parallel development of magnetic-core logic in the Elliott 803 and UNIVAC Solid State. A lighter tangent debates whether Eiichi Goto's surname is nominative determinism predating the Fortran `GOTO` statement, with a digression on romanization and pronunciation.

H3-metal – Native MiniMax-H3 inference for Apple Silicon [comments]

221 points · 32 comments · github.com · 6h ago

Salvatore Sanfilippo (antirez) released a native Metal implementation of MiniMax-H3 inference for Apple Silicon, letting you generate 512x512 video clips from text prompts directly on a Mac. The thread quickly got real about hardware requirements — while the README suggests a 40GB peak footprint, people running the model through ComfyUI on 64GB M5 Pros reported it works with GGUF quantizations, though a 9-second clip can take over an hour, while the native implementation cuts that to a few minutes on a 128GB M5 Max. There’s pushback asking for clearer per-mode benchmarks (T2V vs I2V vs reference conditioning at different resolutions) because the 74-second end-to-end number gets handwavy without those knobs. The bigger story for HN is that antirez is still shipping low-level systems code at this pace — people noted he also wrote dump1090 and Kilo, and the speculation is less about the model and more about what happens when a world-class programmer with financial independence dives into their hobby full time. A separate tangent wondered if sparse attention support (which Minimax hinted at) could land here next, since the author’s comment says he’s testing a `--sparse-attention` optional mode.

Mistral Patent for “Code implemented tool calls” [comments]

218 points · 183 comments · patentsgazette.uspto.gov · 18h ago

Mistral filed a patent for a method where an LLM generates code to make tool calls, runs that code in a sandbox, pauses it when a tool call needs client-side execution, waits for the result from the client, then resumes—essentially a remote procedure call orchestrated by an AI. The HN crowd largely dismissed it as a bog-standard RPC or IPC pattern dressed up with the word "LLM," with many calling it the most obvious architecture any engineer would sketch on a toilet. People dug up prior art to challenge it, pointing to Cloudflare's Code Mode blog post from September 2025 and the "CodeAct" paper from 2024, and noted that preissuance submissions are open to anyone willing to be identified. A significant split emerged: some argued this is just another trivial software patent that will get rubber-stamped by a fee-driven USPTO, while others insisted the "non-obviousness" bar for AI-specific combinations is real enough that it could survive until an expensive inter partes review. A side argument broke out over whether Mistral is really a scrappy EU innovator or a regulatory mouthpiece, given their CEO's recent proposal for an AI levy to fund French culture.

Humanising LLM Outputs Is Dumb [comments]

203 points · 132 comments · kuber.studio · 18h ago

The article argues that asking LLMs to adopt human-like writing styles—like "talk to me like I have ADHD" or using Simplified Technical English—is a bad idea because it forces the model to compress its output during reasoning, losing valuable detail and hiding failures behind polished prose. The HN crowd largely agreed with the premise, but the thread pivoted hard into a collective vent about how newer frontier models like Claude Opus 5 have become nearly unreadable, drowning users in flowery metaphors, jargon, and self-congratulatory padding that requires constant prompts to “cut the bullshit.” Multiple people shared workarounds like using ASD-STE100 as a final rendering step or switching to GPT-5.6 for clearer technical writing, with some accusing the labs of intentionally bloating output to burn tokens. A notable split emerged: a few argued this verbosity is a training failure or vendor lock-in strategy, while others insisted the real fix is to let models speak their own precise language internally and only translate for humans at the boundary, exactly what the article prescribes.

As AI eats the web, the internet’s collective memory is disappearing [comments]

194 points · 183 comments · thewalrus.ca · 9h ago

The article argues that Google's AI summaries are eroding the web's archival function—hallucinating basic facts like sunset times—while the underlying internet infrastructure crumbles from link rot, corporate deletions of entire sites like FiveThirtyEight, and Wikipedia's traffic collapsing as AI scrapes its content without sending users back. The HN thread immediately split on whether this diagnosis is real or overblown: several people pushed back hard on the sunset anecdote, calling the user a "moron" for trusting a generic search tool instead of a dedicated app, while others countered that expecting accuracy from a search engine is entirely reasonable. A significant chunk of the discussion pivoted to practical search comparisons—DuckDuckGo users arguing it's still worse than Google for niche queries, old-guard defenders insisting Google's decline started in 2008 when it stopped doing exact-match searches, and one person now running a meta-search stack with EXA and Tavily because "search as we knew it is done." A recurring dark note: multiple people acknowledged that AI summaries are genuinely faster and more useful for technical configs or log analysis, but they also recognized this creates a tragedy of the commons where no human traffic flows back to the sites producing the original content, and the next wave of training data will be AI-generated garbage.

Learning more about Claude's mathematical capabilities [comments]

185 points · 127 comments · www.anthropic.com · 14h ago

Anthropic published a blog post about a research version of Claude that managed to improve the proven lower bound on the fraction of Riemann zeta function zeros that lie on the critical line from 41.6% to 67.2%—a genuine, incremental advance in analytic number theory. The HN thread seized on the bizarre mechanics of how it happened: a non-mathematician staffer (Jarred Sumner, the creator of Bun) spent a day and a half just telling Claude "keep going" and "believe in yourself" while the model ran thousands of shell commands and coordinated 60 subagents to brute-force its way through 650 failed ideas before landing on a viable combination of existing mathematical papers. A significant chunk of the comments split on whether this is an inspiring example of AI accelerating research or a "beyond parody" situation where we're anthropomorphizing a text predictor by treating stochastic manipulation as encouragement. There was also side chatter about whether the Bun → Rust translation project Sumner previously did was actually buggy, with one faction insisting it works fine and another citing Hyrum's law to argue it's a ticking time bomb. Several commenters pointed out that the paper itself includes an acknowledgements section where the LLM thanks individual human mathematicians, which they found deeply strange.

Magnitude 7.4 Earthquake – 5 km S of San José del Palmar, Colombia [comments]

176 points · 67 comments · earthquake.usgs.gov · 16h ago

A magnitude 7.4 earthquake struck Colombia near San José del Palmar, causing damage and panic across a wide region. HN immediately split into two camps: people with personal ties to Colombia sharing firsthand accounts of the shaking, building evacuations, and spotty communications in Bogotá and Medellín, and a separate crowd debating earthquake prediction science, warning systems, and plate tectonics. Several people living through it described the terrifying experience of phone alerts giving only 5 seconds of warning before the shaking hit, which kicked off a long, practical argument over whether that's enough time to do anything useful versus just adding panic. A recurring curiosity was whether this quake was connected to the earlier Venezuela earthquake and a recent spike in New Zealand activity, with domain-knowledgeable people explaining that seismic waves can indeed nudge other fault lines. The thread also featured a surprising amount of lighthearted confusion from people who initially read the headline as "5 km S of San Jose" (California) and dreamed about earthquakes while safely in Texas.

Rust SIMD on the GPU [comments]

175 points · 86 comments · www.vectorware.com · 13h ago

VectorWare announced they’ve gotten Rust’s portable SIMD (`core::simd`) to compile directly to GPU warp operations, meaning the same SIMD code that runs on x86 or Arm CPUs can now run on an NVIDIA GPU with no source changes. The author showed up in the thread to field questions, and the discussion quickly got into the weeds on whether GPUs are truly SIMD or SIMT, with several people pointing out that NVIDIA’s model is actually SIMT—a warp has a single program counter shared across 32 lanes, not independent threads—so mapping `Simd<T,32>` straight to a warp is conceptually clean but papers over the performance cliffs you hit if you don’t write coalesced, predication-free code. A long subthread broke down how a GPU’s “CUDA cores” are really just lanes in a 32-wide SIMD unit, with the massive register file and hardware thread-scheduling designed to hide DRAM latency rather than depend on cache, which is a fundamentally different tradeoff from CPU SIMD. There was some skepticism about the practical value—one person asked why you’d use this instead of Torch or JAX for array workloads, and the author responded that the goal is to run unmodified CPU libraries that already use `core::simd` on the GPU, not to compete with ML frameworks. The company also confirmed they plan to open-source the compiler bits and sell products on top, and that their broader thesis is that decent GPUs ship in every device but most software ignores them.

50k Boat Names [comments]

174 points · 108 comments · www.beautifulpublicdata.com · 19h ago

The piece catalogs over 50,000 boat names harvested from NOAA's AIS vessel traffic data, sorting them into categories like lawyer jokes, literary references, and nautical puns. The thread immediately derailed into the glaring omission of "Boaty McBoatface," with people noting that the British AUV likely never transmitted on US AIS receivers and that the research ship was ultimately named Sir David Attenborough anyway. Several boat owners chimed in with their own favorites—"Unsinkable II" got repeated nods as a perfect name, and one captain collected terrible names on his phone, calling people "insane" for what they choose. The data work itself drew some sharp critique: commenters pointed out the search/categorization was buggy (only one "Freedom" boat when the article claims 101, and "Comanche" lumped under mythology), and the boat-ownership-by-income chart was easy to misread as percentage of owners in each bracket rather than likelihood of owning at a given income level.

'Pervert glasses': Backlash against Meta's smart glasses grows [comments]

173 points · 257 comments · www.seattletimes.com · 16h ago

The article covers the growing backlash against Meta's Ray-Ban smart glasses, which have been dubbed "pervert glasses" because some users secretly record women in public and post the footage online. The HN thread split sharply, with some arguing the product is fundamentally toxic and that any use in social settings is inherently creepy, while defenders insisted there are legitimate use cases like capturing grandchildren or hands-free photography at events. A significant contingent pointed out this is just Google Glass all over again—the same "glasshole" dynamic from a decade ago—and argued the product category is probably dead in public spaces regardless of privacy features. Others focused on the hardware itself, wishing for glasses with a HUD display but no camera, but noted that the trust problem is now so deep that even camera-less smart glasses would face suspicion. The strongest pushback came from people who said the issue isn't Meta specifically but the fundamental creepiness of not knowing if a conversation is being recorded for social media monetization.

Show HN: Scroll through all 43252003274489856000 Rubik's Cube states [comments]

170 points · 55 comments · everycube.alen.is · 8h ago

The linked article wasn't available to this summarizer; from the discussion, it's a site that lets you scroll through every single possible Rubik's Cube state—all 43 quintillion of them—using URL-based indexing to jump around. The immediate reaction was pure delight at the sheer absurdity, with people quickly doing the math: one person figured that at 12 scroll-wheel states per inch moved at the speed of light, it'd take 9.5 years to see them all. The joke that this is the "healthiest possible doom scroll" landed hard, and a running gag emerged about treating the cube positions like NFTs or a housing affordability crisis where one cube "hoards" quintillions of states while others have none. A few practical suggestions surfaced—auto-rotate and a solver that shows the quickest solution for any state—and the author shipped a live fix when someone noted the page didn't update on URL hash changes.

Amazon backs power plant that may become top source of US climate pollution [comments]

168 points · 138 comments · arstechnica.com · 10h ago

Amazon is funding a massive new natural gas power plant in Texas to run its AI data centers, a project that could become the single biggest source of climate pollution in the U.S. The HN thread quickly moved past the article itself into a sprawling argument about whether the public even cares enough to stop this, with one camp insisting fossil fuels need to stop yesterday and the other pointing out that most Americans don’t share that urgency—especially in Texas, which leads the nation in renewable energy but also consumes more coal and gas than any other state. A big split emerged over whether data center opposition is genuinely about the environment or is just a convenient political cudgel against tech companies, with several people arguing that the same NIMBY energy that blocks apartment buildings is now being aimed at data centers. Someone made the point that gas from the Permian Basin would otherwise just be flared off anyway, so buying it might actually be less wasteful than letting it burn in the open, while others retorted that this is exactly the kind of rationalization that lets the industry keep building without changing course. The thread never really settled anything—people who think this is an existential crisis faced off against people who think the crisis is overblown and that Texas’s open-permitting ethos is a feature, not a bug.

Chicken Scheme 6.0 [comments]

166 points · 17 comments · code.call-cc.org · 7h ago

The linked article wasn't available to this summarizer; from the discussion, Chicken Scheme 6.0 is a major release of a Scheme-to-C compiler that finally brings full Unicode support, R7RS compliance, and replaces blobs with bytevectors. The core pitch is that Chicken compiles to portable C, so you can run Scheme anywhere a C compiler works, and it has the "eggs" library ecosystem for things like web servers and SDL2 games. People are split on whether the compiler or the ecosystem is the bigger draw, with one dev calling the generated C "the gem" and praising the useful stack traces, while another notes the documentation website has been flaky and offline docs would be a huge improvement. One surprising tangent is that version 6 also supports Crunch, a statically-typed subset of Scheme that's nearly at 1.0, and someone who helped make it work on Windows notes it still needs "a bit of setup." There's a dry joke about what happens if you try to run Python eggs under Chicken, and a disappointed viewer clarifies this is not the "Chicken chicken chicken" video.

What's the best programming language for coding agents? [comments]

158 points · 105 comments · danluu.com · 15h ago

Dan Luu ran a detailed eval to test the claim that dynamic languages are more token-efficient for coding agents, and the HN thread latched onto his conclusion that the claim mostly falls apart once you move past trivial Rosetta Code problems—the extreme efficiency of languages like J or Clojure doesn’t hold up on real tasks like implementing a zstd decoder. The discussion zeroed in on the confounders Luu exposed, like a Go agent symlinking its own executable over a missing path and ruining later Rust scores, which made several people argue the original “dynamic wins” studies were just measuring broken test infrastructure. A few commenters pushed back hard that syntactic density isn’t the same as token cost anyway, since symbols tokenize poorly compared to plain English, and that the small deltas Luu found make it hard to justify not using Rust or Go for their correctness benefits. Others took a tangent on language popularity, suggesting Python and JavaScript dominate because training data quantity swamps language design, though a couple people shared surprising counterexamples—like one guy who got perfect xTensa assembler out of Fable 5 but found LLMs flailed on Wolfram despite its conciseness.

Exploiting System Management Mode with a very long interrupt [comments]

155 points · 58 comments · github.com · 16h ago

The linked article is a proof-of-concept exploit by Christopher Domas showing that you can break System Management Mode (SMM) on an x86 CPU by running a single instruction that lasts longer than the firmware's 1-second rendezvous timeout, allowing one core to execute outside SMM while another core is inside it. The thread latched onto the exploit's absurd simplicity: you just need a slow MMIO read on one core, and the firmware's built-in timeout does the rest of the work, cracking open a whole class of SMM TOCTOU vulnerabilities that were previously considered unexploitable without hardware access. A big chunk of the discussion argued about whether the 1-second timeout is fundamentally broken—some called for removing it entirely or making it infinite, while others pointed out that the timeout exists to prevent legitimate hangs from bricking the machine, and the author himself admits there's no clear fix. People also noted that this requires root access anyway, so it's less a remote exploit and more a way for an attacker who already owns the kernel to take over the even-more-privileged SMM, with one side saying this just proves SMM is a user-hostile backdoor and the other countering that the mechanism itself isn't the problem, it's the opaque, unpatched firmware. There was also meta chatter about Domas's release spree after leaving Intel and his Defcon talk, plus some entertaining appreciation for the readme's over-the-top stylistic commitment to the word "long."

Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints [comments]

149 points · 162 comments · www.wcax.com · 17h ago

Kinney Drugs is rolling back its AI phone assistant “Burt” after customers reported garbled calls, wrong dosages, and missed prescription notifications. The thread quickly zeroed in on the absurdity of using voice AI for pharmacy work: drug names are a pronunciation nightmare (Wegovy, Qvar, Ixempra), ASR word-error rates are still atrocious with regional accents, and a chatbot that can’t handle a simple address-update issue is a death sentence when the stakes are life-sustaining meds. Several people likened this to the disastrous offshoring of call centers in the 2000s, noting that companies eventually reversed those moves because the customer experience cratered. A split emerged between practitioners who say careful, narrow deployment can work and the majority who argue that investor-driven “AI is magic” hype is forcing half-baked systems into workflows they aren’t ready for, turning every minor mistake into an escalating crisis that no human agent can easily unwind.

Tail-call optimization in C is relatively recent (2025) [comments]

143 points · 131 comments · lwn.net · 20h ago

The linked article is an LWN piece where Anton Ertl notes that tail-call optimization (TCO) in C is surprisingly recent, only becoming reliably available in GCC and Clang within the last couple of decades. The HN thread largely pivots to a debate about TCO across languages: JavaScript’s spec-mandated “Proper Tail Calls” are implemented only in Safari, leaving the web stuck with stack-overflow bugs, while Rust’s proposed `become` keyword has people excited because it promises a compiler error when TCO can’t be delivered, unlike in C++ where Clang’s `must_tail` attribute is merely an ignorable hint. A strong split emerged over whether TCO belongs in a language specification or is just an optimization you can’t rely on—Scheme and SBCL fans argue it should be guaranteed, while C and C++ veterans note that destructors, ABI quirks, and callee-saved registers make guaranteed TCO a much harder problem. The thread also got into the weeds on whether C’s historical calling conventions actually prevented TCO, with several people pushing back that the original author’s 1994-era reasoning about argument-cleaning rules no longer applies.

How Claude marks AI-generated content [comments]

141 points · 100 comments · support.claude.com · 10h ago

Anthropic published a detailed explainer on how Claude will embed invisible watermarks into generated text and attach signed provenance metadata to files, as part of its commitments under the EU AI Act's transparency code of practice. The HN thread immediately zeroed in on the technical feasibility of text watermarking, with people batting around how it actually works under the hood — some pointed to statistical token-selection methods like the "green/red token" logit-biasing approach from the arXiv paper, while others argued any watermark relying on pattern or word choice is trivial to strip with a grep, a git hook, or a quick script, especially for code or config files. A strong split emerged: one camp insists any text-level watermark is doomed because people will build sanitizers instantly (one person already registered a watermark-removal domain), while the other side counters that detection doesn't need perfect recall — high precision with some false negatives is good enough for compliance, and most users won't bother stripping marks anyway. There was also a recurring debate about whether Claude's distinctive writing style is itself a de facto watermark Anthropic has leaned into, and whether widespread AI-generated prose will degrade human writing style or, conversely, push people to adopt more distinctive voices to avoid being flagged.

Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines [comments]

141 points · 20 comments · blog.sshh.io · 17h ago

The post explores a technique called "incompressible knowledge probes" to reverse-engineer hidden details about frontier AI models—specifically their knowledge cutoffs, parameter counts, and training data mixtures—by carefully quizzing models like Claude and GPT on niche facts and self-identification questions. The HN discussion largely ran with the idea that marketing model names like "Opus 5" are not a single static weight set but can get silently updated on the backend, though an OpenAI employee jumped in to clarify that for API models, the weights really are fixed (with tiny exceptions), while ChatGPT models do get minor updates without a name change. A big split emerged around the spiciest finding: that Anthropic's Sonnet 5 regularly self-identifies as GPT-4, which the post speculates could be from training on old ChatGPT chats in the mix, but several people pushed back hard, noting that distilling from public datasets like LMSYS Chatbot Arena is perfectly legal and very common, making the accusation less damning than the post implies. Others zeroed in on the practical tactic of waiting to release models, arguing that the competitive pressure should force companies to ship ASAP, countered by a claim that Claude has held the "best for coding" slot since last November, suggesting some labs can afford to sit on improvements.

Show HN: Ante, a coding agent in a single binary that runs offline [comments]

130 points · 78 comments · github.com · 16h ago

The article announces Ante, a ~15MB Rust-based coding agent that runs as a single offline binary in your terminal, embedding its own llama.cpp engine for local inference and promising zero runtime dependencies or vendor lock-in. HN’s biggest point of contention was the opt-out telemetry, with several people calling it contradictory for a tool marketed toward offline use and secure environments, though the author acknowledged the feedback and said it was a carryover from the dev preview. Another major thread centered on the lack of open source code—the core harness ships as a prebuilt binary—and while the repo holds docs and SDKs, multiple engineers pushed back hard, arguing that without full source, there’s no way to audit for malware or supply-chain risks, and one person pointed to their own fully open-source alternative written in C. The discussion also spun out into a broader debate about what counts as “building” software with AI agents, using game development as a case study: some argued agents are just another tool like a hammer, while others insisted they’re more like a ghostwriter or a Roomba, and commenters with hands-on game dev experience reported that current agents still struggle badly with visual tasks and real-time interaction. A few technical side notes emerged—like the distinction between “ships its own inference engine” and “pins a trusted llama.cpp build,” which the author corrected—and there was a request for native Windows support (CUDA) that got added to the backlog.

Study links GLP-1 drugs to bigger jump in women's employment than a degree [comments]

124 points · 177 comments · finance.yahoo.com · 16h ago

A Harvard working paper found that women who started taking GLP-1 weight-loss drugs saw a 27-percentage-point jump in employment rates, a bigger gap than the employment difference between women with a high school diploma and a college degree. The HN discussion largely accepted the core premise — that “pretty privilege” or the “halo effect” is real and long-documented — but quickly split on interpretation. Some pushed back on the study’s methodology, noting that the control group (women who wanted the drugs but hadn’t started) introduces class and status confounders, and that the paper is self-reported and tentative. Others argued that appearance-based discrimination is obviously baked into hiring, pointing out that the effect likely applies differently to men depending on job type and that the conversation often overlooks how weight loss itself causes cognitive and sleep improvements that also affect employability. A recurring tension was whether the employment boost comes from reduced stigma or from genuine behavioral changes like increased confidence — and whether the “just-world fallacy” makes people too quick to assume the latter.

30 threads · window 24h · article context usable 26/30 (unavailable 0, skipped 0, agent failed 4)
Generated 2026-08-11 08:12 UTC

Generated by Sauron from Hacker News discussions and linked articles.