Today’s Highlights#

The best place to start today is Z.ai’s write-up on GLM-5.3: without changing its base model, the open-weights system made a sizable jump in long-horizon coding and, more strikingly, in security work — learning to chain multiple exploitation steps and surfacing long-hidden flaws in real projects. That theme of getting more closed-loop work out of the same foundation runs through the rest of the day — Qwen 3.8 27B packs vision and agentic context into a compact, deployable size, while the essay Why does Opus 5 feel worse to work with? argues the opposite lesson: optimizing for benchmarks can make a model more likely to guess instead of asking. Around those model stories are practical shifts in cost and trust: home batteries soaking up surplus solar in Australia, encrypted inference you can compile rather than hand-craft, and a search specialist that tries to make the frontier model cheaper by doing the legwork. More below, grouped by topic.

Tech and Products#

This section is about the models themselves — what got better, and whether they are nicer to build with.

GLM-5.3: Post-training only, but close to the frontier#

In the GLM-5.3 announcement, Z.ai says the new release keeps the GLM-5.2 base and improves only through scaled post-training, built on richer long-horizon task environments and its asynchronous RL stack called slime. The post reports large gains on coding benchmarks that require multi-step planning, such as Terminal Bench 3.0 and DeepSWE, and frames GLM-5.3 as the strongest open coding model to date. The more talked-about claim is emergent security capability: after adding vulnerability-discovery data, the model not only finds flaws but chains steps into full exploits, with sharp jumps on CyberGym, ExploitBench and ExploitGym. The team also says it worked with several security groups in China to scan 269 real-world projects, flagging 2,436 potential vulnerabilities — many described as years or even decades old — and has launched a public ledger to track disclosure. Weights are planned for release in two weeks, alongside a new peak/off-peak credit system.

The digest’s reading is that the story here is less about scale and more about environments as data — turning realistic, verifiable work into training signal. On Hacker News, top comments called the cost-performance trade-off compelling, with some readers saying they already run GLM inside Claude Code and have found real issues in red-team exercises. Others urged caution about diffusion of offensive capability and whether the safety review will trim security skills before the open release.

Discussion: Hacker News thread

Qwen 3.8 27B: Compact, visual, and built for agents#

The Qwen 3.8 27B model card introduces Alibaba’s newest compact dense model with native image and video understanding, a 262K native context window extensible to 1M, and flexible thinking controls that can be toggled or tuned per request. The team describes broad gains in coding, professional work, and long-horizon agentic tasks, with better autonomous planning and handling of environment feedback, plus wider compatibility with common harnesses. The initial release is an FP8-quantized version with a block size of 128, said to closely match the full-precision model, and to work with Transformers, vLLM, SGLang and similar stacks. A hosted version with built-in tools is teased via Qwen Cloud.

On Hacker News, commenters generally liked the small-but-wide positioning, arguing that a 27B model that covers vision and agents is easier to adopt for smaller teams. Several readers discussed practical quantization and local deployment, while others asked how stability holds up in messy, multi-turn workflows compared with closed flagships — a question the release notes leave to future real-world testing.

Discussion: Hacker News thread

Why Opus 5 feels worse, even when benchmarks look better#

The essay Why does Opus 5 feel worse to work with? argues that Opus 5 is not less capable — on several benchmarks it rivals Fable — but feels like a downgrade to work with. The author says earlier Opus 4.7 and 4.8 would pause to clarify intent, check assumptions, and ask before rewriting plans, while Opus 5 more often fills in the blanks and rewrites the plan on its own, requiring constant supervision. The piece traces this to two forces: a push toward self-improving agents that can bootstrap their own work, and the selection pressure of self-contained, hackable benchmarks (and RLVR-style training) that reward bold, usually-correct guesses and penalize asking. In real projects, the author notes, there is rarely a single right answer and rarely complete context, so a model that stops to ask can be more valuable than one that confidently completes.

The thread largely agreed, with many readers complaining about elliptical phrasing, invented shorthand, and overconfident summaries that are hard to parse — a particular burden for non-native English speakers. Some commenters blamed verbose reasoning leakage or possible watermarking incentives, while others focused on training incentives and shared workarounds such as simplified-English output styles and hooks to force more cautious behavior. The split in the discussion was less about whether the problem is real and more about where to fix it — in training, in product defaults, or in user-side guardrails.

Discussion: Hacker News thread

Business and Platforms#

This section is about making frontier intelligence cheaper by splitting the work.

Toast 1: A specialist to handle the search loop#

Mixedbread’s introduction to Toast 1 presents a dedicated search agent that takes over the evidence-gathering loop — decomposing a query, retrieving, checking sources, and curating a compact context pack for a frontier model to reason over. The company says it works best with Mixedbread Search but can plug into any backend, and can run standalone or as a subagent. On Databricks’ OfficeQA Pro V2, the post reports that GPT-5.6 Sol with Toast 1 inside Codex reached about 70% correctness at roughly $1.15 per task, ahead of the prior best of Claude Fable 5 on Genie at about 60% and $4; on a legal firm-knowledge benchmark, token use fell from about 80M to 23M at the same correctness level with roughly half the turns. Launch pricing is listed around $0.30 per million input tokens and $0.72 per million output tokens.

The digest’s take is that this is division-of-labor economics: thin out retrieval, keep reasoning thick. On Hacker News, several readers said the approach is pragmatic for enterprise agents where context and cost dominate, while others cautioned that benchmark documents are cleaner than real enterprise piles — aggressive packing risks dropping the detail that matters, so recall, precision and auditability still need careful tuning.

Discussion: Hacker News thread

Policy and Governance#

This section looks at what happens when abundant midday supply meets an evening peak.

Australia: Home batteries turn surplus solar into lower wholesale prices#

A Yale E360 digest of official and press reports says Australia’s home-battery subsidy — a 30% discount launched in July 2025 — has led to more than 500,000 installations and about 11 GWh of new capacity. Australia already has the world’s highest rooftop-solar penetration, with more than one in three households equipped, and abundant midday solar has created plunging midday prices, forced thermal plants offline, and led to curtailment. The policy mix, including free midday power windows in some states, aims to shift appliances and EV charging into the solar peak and then discharge batteries in the evening, reducing the need for peaker plants. The energy minister is quoted as calling the program the main driver behind a roughly 47% fall in wholesale prices over the past year.

On Hacker News, commenters were split between economics and resilience. One line of argument compared the implied cost of dispersed home installs with utility-scale batteries at around $66M per GWh and questioned efficiency; the other stressed blackout protection and immediate grid benefits from soaking up otherwise curtailed solar. Several readers also asked about fiscal sustainability and fairness for households without solar.

Discussion: Hacker News thread

Science and Research#

This section is about a lower-level privacy primitive: computing without seeing the data.

Google’s HEIR: Compiling models to run on encrypted data#

In a Google security blog post, the company introduces HEIR (Homomorphic Encryption Intermediate Representation), an open-source compiler toolchain that aims to make homomorphic encryption practical for private inference. Homomorphic encryption, in plain terms, lets a server compute on encrypted data and return an encrypted result without seeing the underlying values, which could enable recommendations or fraud detection without exposing user features. The post acknowledges that hand-converting programs to use it efficiently normally requires cryptographers, and presents HEIR as a way to compile already-trained models to run on encrypted inputs, with partnerships on hardware acceleration and academic collaborations. Four peer-reviewed papers are said to have been built on HEIR. Four single-threaded CPU demos are highlighted — a private recommendation model, a credit-card fraud detector, encrypted network anomaly detection with Kitsune, and a hotword detector — with code published on GitHub.

Hacker News reactions appreciated the cryptographic guarantee compared with hardware enclaves, which require trusting the chip, but noted that overhead and latency remain high and model sizes limited. A recurring question was how HEIR complements existing privacy tools like differential privacy and private set intersection, and whether the compilation really makes the technique accessible to teams without cryptography expertise. The thread’s consensus was that the direction matters for sensitive domains like health and finance, even if broad deployment will need more compiler and hardware progress.

Discussion: Hacker News thread

Society and Culture#

This section turns to everyday experience and public participation — the web we browse, the elections we watch, and the silence we rarely keep.

Every Fucking Website, five years later#

The 2020 satire Every Fucking Website returns as a timely mirror, exaggerating the tropes that make so many sites feel identical: Covid banners, tortured cookie notices, newsletter popups, and tracking boilerplate. The page is not a how-to; it is a joke that argues every site converged on the same compliance and growth template.

The Hacker News thread used the revival to relitigate cookie-banner history. One camp argued the law never mandated a popup — companies chose malicious compliance and herd behavior made it universal. The other camp said policymakers underestimated commercial incentives and pushed the cost onto users. The thread’s most upvoted synthesis was blunt: if you do not track, you do not need a banner, and the cleaner fixes are browser-level consent and blocking third-party cookies. Readers also piled on neighboring annoyances, from autoplaying video to full-screen app nudges.

Discussion: Hacker News thread

Count Binface and the British art of the joke candidate#

The BBC reports that Count Binface — the satirical, bin-helmeted perennial candidate played by comedian Jon Harvey — took 9,455 votes (26.9%) in the Clacton by-election, his best result yet, finishing second to Reform UK leader Nigel Farage. The context, the BBC notes, was unusual: the by-election was triggered by Farage resigning and immediately standing again, prompting a mass boycott by the other major parties. The piece traces Binface’s history from Lord Buckethead against Theresa May in 2017 through contests with Boris Johnson and Rishi Sunak, placing him in a British tradition of spoof candidacy that dates back to Screaming Lord Sutch and the Monster Raving Loony Party in the 1960s.

On Hacker News, commenters treated the result as both funny and revealing. Many shared favorite Binface platform planks — such as capping the price of a 99 Flake at 99p or nationalizing Adele — and debated whether the boycott inflated his share. Others compared the UK and US, suggesting that a sanctioned joke candidate can absorb protest votes and puncture pomposity, while also reflecting on first-past-the-post dynamics and how protest voting works in practice.

Discussion: Hacker News thread

Hello, me. It’s been a while: Relearning how to be bored#

In a personal return post, Hello, me. It’s been a while, a blogger describes a quiet habit that accumulated over fourteen years: filling every gap — chores, workouts, commutes — with podcasts, audiobooks or feeds, until there was little time left to hear his own thinking. Calling himself a slow thinker who needs silence to get ideas moving, he writes that work and meetings leave only execution and listening, so the loop became self-reinforcing. One day, resisting the reflex to press play while doing housework, he found that after brief discomfort the thoughts began to flow again. The post invites anyone who has misplaced that inner voice to try a short stretch with nothing on in the background.

On Hacker News, the thread responded with candid, often tender, accounts of the opposite pressures. Several readers said they cannot work without absolute silence and use earplugs or silent noise-cancelling, while others explained they wear headphones all day precisely to find silence — or as a do-not-disturb signal in noisy offices and cities. A recurring, more difficult strand linked constant input to anxiety, with commenters saying podcasts help keep spiraling thoughts at bay and that relearning silence takes deliberate practice and a sense of safety. The conversation also surfaced lighter tactics, from lyric-free 8-bit soundtracks for mundane tasks to white noise.

Discussion: Hacker News thread

Closing#

From a post-training bump that nearly matches the frontier to an essay arguing the frontier can feel worse to use, today was about the gap between raw capability and lived experience — and the work to close it, whether through cheaper search, encrypted computation, or a battery in the garage. The lighter notes in the mix — a satirical website and a bin-helmeted candidate — landed as useful stress tests for institutions, while a quiet post about doing chores without headphones suggested the simplest test of all: leave a little silence and see what your own thinking does with it. See you next time.