15 STORIES

Daily Byte.

A byte-size daily digest for engineers who value depth over noise.
Wednesday, August 26, 2026 · Edition 2026-08-26
Engineering
Engineering

Independent voices from engineering blogs — what shipped, what hurt, and what language designers think matters this week.

ENGINEERING

How to evaluate LLMs before production

LLMs can perform well on clean benchmarks yet struggle with the ambiguous, inconsistent inputs they encounter in production. The authors recommend starting with the product decision rather than the model, defining which mistakes are acceptable and which guardrails must remain within thresholds, treating precision and recall as non-interchangeable metrics. Offline evaluation should function like integration testing, rerun whenever prompts, models, or pipeline logic change, with results recorded so each run can be compared against a known baseline. Experiments should change one major variable at a time to keep causation clear, and prompts and configurations should be versioned like code. Finally, offline evaluations must closely mirror production conditions, preserving the surrounding context, formatting, and ambiguity of real inputs, since differences between curated datasets and production data can materially skew results.
GitHub Blog
ENGINEERING

Your alt text passes automated checks. That doesn’t mean it’s any good.

We built a plugin for the GitHub Accessibility Scanner to make sure your alt text is actually accessible. Here's how it works. The post Your alt text passes automated checks. That doesn’t mean it’s any good. appeared first on The GitHub Blog.
GitHub Blog
ENGINEERING

10x Engineer DNA

Arpit Bhayani, who focuses on engineering, databases, and systems, argues that 10x engineers earn their reputation not by building complex systems but by consistently finding simple, stable, and manageable solutions to complex problems. The article, titled "10x Engineer DNA," presents this philosophy as the pathway to becoming a 10x engineer. Bhayani currently works as a Principal Engineer II at Razorpay, building Agent Studio, and previously held roles including staff engineer at GCP Memorystore and Dataproc and Director of Engineering at Unacademy. He is also the creator of DiceDB and has worked at Amazon Fast Data. Beyond his engineering career, he produces engineering videos on YouTube, teaches courses, and writes a newsletter read by 145,000 engineers covering system design, distributed systems, and clever algorithms.
Arpit Bhayani
ENGINEERING

1D Terrain Generation

Generating pseudorandom one-dimensional terrain for games starts with a naive approach that assigns random heights to each point along the X-axis, but this produces abrupt spikes that don't resemble real-world terrain. Applying interpolation methods, either Linear or Cosine, between sampled points smooths transitions, with Cosine Interpolation producing noticeably smoother curves than Linear by plotting intermediate points on a cosine curve rather than assuming collinearity. To further improve realism, the technique of superposition combines multiple sampled terrains with different sampling frequencies through a normalized weighted sum, where weights increase as sampling frequency decreases. Typically six terrains are combined, with the most densely sampled terrain receiving the lowest weight (0.03125) and the least sampled receiving the highest (1), giving the final terrain both smoothness and subtle variations. This superposition-based method closely resembles Perlin Noise, the multi-dimensional terrain algorithm for which Ken Perlin received an Academy Award for Technical Achievement.
Arpit Bhayani
ENGINEERING

6 Simple Strategies to Cracking Any Tech Interview

Arpit Bhayani shares six strategies he used throughout his career to crack tech interviews. He recommends using Python during interviews because its concise syntax allows candidates to express solutions more quickly than writing verbose Java or C++ code, though he notes some interviews require a specific language and candidates should clarify this with the interviewer first. Bhayani is a Principal Engineer II at Razorpay, where he is building Agent Studio, and his background includes ex-staff engineering roles at GCP Memorystore and Dataproc, founding DiceDB, work at Amazon Fast Data, and serving as Director of Engineering for SRE and Data Engineering at Unacademy. Beyond interviews, he creates engineering content through YouTube videos and offers courses such as the Applied AI Masterclass, System Design Masterclass, and System Design for Beginners, all listed under Relog Deeptech Pvt. Ltd. His newsletter, Arpit's Newsletter, is read by 145,000 engineers on LinkedIn and Substack and covers topics like system design and distributed systems.
Arpit Bhayani
Papers

Recent arXiv papers in AI, ML, and programming languages — chosen for working engineers, not just benchmark wins.

PAPERS

How to Train a Critic Stably and Efficiently

The paper "How to Train a Critic Stably and Efficiently" by Penghui Qi, Xiangxin Zhou, and Wee Sun Lee addresses instability in critic-based reinforcement learning for large language models. While group-based methods like GRPO avoid training a critic by sampling multiple responses per prompt, standard critic-based training recipes that estimate token-level advantages from a single response tend to be unstable. To address this, the authors develop Best-Practice Critic Optimization (BPCO), which combines DPPO, reward-bounded value predictions, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can condition it on reward-defining information such as reference answers or grading rubrics that remain hidden from the policy. Experiments on mathematical reasoning tasks with models ranging from 1.5B parameters to 30B-A3B mixtures of experts show that BPCO consistently improves a critic-based baseline and matches or exceeds a group-based baseline while sampling just one response per prompt.
arXiv · cs.AI
PAPERS

ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld is an interactive world model designed to follow user actions, maintain long-horizon memory, and stream video in real time. It separates short-horizon control from unbounded memory during training by using mixed per-head attention windows, where most heads attend to recent frames while a small set of global heads attends over the full history, supported by random head routing and chunk dropping. At inference, a bounded KV cache backed by a pose-indexed landmark bank retrieves the nearest landmarks to the current pose, keeping the entire past under a fixed budget. A data engine aligns eight diverse video sources on one physical action scale, and palindrome trajectories supply the revisit evidence needed for memory training. Distribution-matching distillation within a LoRA adapter compresses sampling to four steps, enabling 704x1280 streaming across photorealistic, game-style, and stylized worlds. Evaluated against six recent interactive world models under a three-axis protocol, ReWorld achieves the best control fidelity, including 11.95° rotation error, and regenerates starting views on 64-second out-and-back rollouts where sliding windows and full attention fail.
arXiv · cs.AI
PAPERS

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench is a benchmark of 20 whole-repository migrations covering four kinds of technical debt, designed to test whether coding agents can autonomously perform long-horizon stack migrations. Existing benchmarks suffer from "Blindness," where agents copy the original implementation to pass behavioural tests without performing the actual migration. To address this, the authors propose a three-stage evaluation protocol: Migration Audit verifies the migration occurred, Behavioural Tests measure correctness with a fixed test suite, and Agentic Verification uses six independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 5.4% pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model, claude-opus-5, scores 47.0 out of 100. The results reveal that migration completeness and behavioural correctness are distinct abilities, and capability varies significantly across migration categories, with agents scoring 31.4 on build toolchain rewrites but only 5.6 on language rewrites.
arXiv · cs.AI
PAPERS

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, auditor shortages, and costly field inspections. The authors propose Expert-Grounded Distillation (EGD), a framework that transfers institutional road safety expertise into a compact vision-language model. EGD features a quantified expert-grounding stage where a teacher model is calibrated against authoritative field audits, achieving substantial agreement with expert risk assessments (Cohen's kappa = 0.74). The calibrated teacher generates structured supervision distilled into an 8-billion-parameter student vision-language model using Low-Rank Adaptation. The authors also introduce BD-ARSA, an open Bangladeshi visual road safety audit dataset of 21,947 image-audit records with near-national coverage, and EG-ARSA, the first vision-language model for this task. Results show grounded fine-tuning improves ordinal risk assessment over the zero-shot baseline, while blind expert evaluation reveals the compact student outperforms both its 31-billion-parameter teacher and Gemini-2.5-Flash.
arXiv · cs.AI
PAPERS

Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography

The paper introduces Phy-BP, a non-invasive blood pressure estimation framework based on triaxial bodyseismography (BSG), which extends traditional ballistocardiography to address misaligned representations caused by body-bed interaction variations and individual hemodynamic changes. An adaptive quality-control algorithm selects BSG segments enriched with cardiogenic components by jointly evaluating neighboring beat patterns and universal cardiogenic templates, dynamically filtering out low-quality measurements. A physical model describing 3D wave propagation within the body-bed system is embedded into the deep learning architecture to characterize the intrinsic coupling among the three BSG axes driven by a single cardiogenic excitation, thereby aligning multi-axis features during training. Experiments on a 162-hour hospital dataset collected from 21 subjects demonstrate that the physics-constrained deep learning model improves robustness against real-world distortions and enhances generalizability, particularly when training samples are limited.
arXiv · cs.AI
Tools

Hacker News projects worth your attention — open-source releases that solve a real problem.

TOOLS

Nitter and XCancel receive cease and desist notices

The zedeus/nitter GitHub repository was archived by its owner on August 25, 2026, becoming read-only following reports that Nitter and XCancel received cease and desist notices. On the same day, an issue titled "All public nitter instances do not work." was opened, reporting that every public instance displayed the same "Instance has been rate limited." error message. The issue, posted by user AlexandrPutenikhin, included a link to an archived version of the twiiit.com instance as supporting evidence of the problem. The post quickly gained attention, accumulating 28 comments and 8 thumbs-up reactions from the community. The repository, which had amassed 13.5k stars and 824 forks, is no longer accepting contributions or updates as a result of the shutdown.
Hacker News621 pts
TOOLS

Show HN: LatticeDB – Like SQLite but for graph databases

LatticeDB is an embedded, single-file property-graph database written in Zig that combines graph traversal, HNSW vector similarity search, and BM25 full-text indexing within a unified query layer. Designed for local-first use on a single machine, the entire database lives in one portable file with no server or configuration required. It supports Cypher-style queries and offers performance benchmarks of 0.13 μs for node lookups and 0.83 ms for vector search across 1 million vectors while maintaining 100% recall. The system shares a single transaction and WAL path across graph writes, durable named streams, and a built-in graph changefeed. It is positioned for relationship-heavy workloads such as Graph RAG, agent memory, and local knowledge tools, with the repository showing 220 stars, 8 forks, and 436 commits at the time of the Show HN post.
Hacker News106 pts
TOOLS

Show HN: I made a Raspberry with Qwen my local car AI

CarWatch runs as a fully offline car agent on a Raspberry Pi 5 (16 GB, ~300 €) hosting a Qwen3.6-35B-A3B model locally, achieving 3.5 tok/s generation and 25+ tok/s prompt processing at 65 °C sustained. It joins GroupMind chat rooms as @gle, sending messages about departures, arrivals, trip summaries, and dashcam clips when impacts occur, with approvals routed through CodeWatch. The system answers owner's manual questions via lexical RAG over the car's own 745-page manual with page citations, refusing to answer outside the manual's scope. Hands-free voice input runs entirely on-Pi through energy VAD and whisper.cpp, while grounded self-knowledge reads live hardware metrics like temperature, throttling, fan, memory, disk, and network status. systemd services auto-start the entire stack on boot, and the car pulls its own updates over a three-tier connectivity fallback spanning phone hotspot, home WiFi, or its own access point.
Hacker News98 pts
TOOLS

HelloAssembly: The smallest possible complete Windows application (2021)

HelloAssembly is a GitHub project by PlummersSoftwareLLC dedicated to creating the smallest possible complete Windows application without compression. It was inspired by two episodes of Dave's Garage on YouTube, including "Hello, Assembly! Retrocoding the World's Smallest Windows App in x86 ASM" and a follow-up comparing C versus assembly. The repository contains multiple versions of the application, including an optimized original called TinyOriginal, a Lasse version that applies shell coding tactics, and a Theron version using a manually written PE header, while a QRCode directory contains dead code preserved for posterity. The goal is to produce an executable that runs a Windows message loop, features a title bar with working minimize, maximize, and close buttons, has a system menu, and paints a background with centered text reading "Dave's Tiny App." The project notes that the aggressive optimization techniques used may cause virus scanners to flag the resulting executables as suspected malware, potentially requiring users to whitelist them.
Hacker News83 pts
Discussions

What the HN and Lobsters communities are debating today — threads with substance behind the noise.

DISCUSSIONS

I cannot survive from burnout

The original poster describes experiencing burnout two years ago, followed by divorce and relocation, leaving them living alone with seven cats and carrying significant debt. Despite delivering high-quality client work, they can only manage 5–6 hours weekly on paid projects while easily focusing on personal side projects. Commenters suggest the pattern of pursuing novelty over obligation may indicate executive function issues, and several recommend an ADHD evaluation. Multiple users share personal experiences where ADHD diagnosis and medication dramatically improved their ability to focus on required tasks. The thread also debates whether therapy should aim for a defined endpoint or function as ongoing maintenance, with some favoring goal-oriented approaches and others viewing therapy like regular exercise. The poster expresses gratitude for the unexpected volume of supportive responses and is seeking a new therapist.
Lobsters