I write about CS, ML/AI, and software engineering.
UDP connect() Sends Nothing, and getsockname() Gives You Your Local IP
A machine with Wi-Fi, a VPN and a container bridge has several addresses, and a hostname lookup can return the wrong one. Connecting a UDP socket to a destination makes the kernel pick the source address it would use, and getsockname() reads it back without a single packet being sent. Checked on Linux with three interfaces: five destinations, five answers, no IPv4 or ARP frames on the wire.
Socket Activation Keeps Connections Waiting, Not Refused, During a systemd Restart
A runbook for replacing a running daemon without dropping requests, tested on a small TCP server under real systemd. A plain restart refused or reset connections every time; starting the new version beside the old one avoided refusals but still reset a connection in some runs, and a kernel setting made those resets go away; socket activation produced no errors. Also covered: draining in-flight work against TimeoutStopSec, and a release swap that checks the new version and rolls back.
Expo's Fingerprint Runtime Version Changes When You Edit extra, Not Your JavaScript
An over-the-air update is JavaScript that calls into a native binary you cannot change. Treat the runtime version as the ABI version of that binary. Experiments with @expo/fingerprint 0.20.13 show which edits change the hash (extra, version, build numbers, native modules) and which do not (JavaScript, pure-JS dependencies), and a skip list that silently drops a default.
systemd Kills the Updater Your Service Started, Even With setsid
A daemon that updates itself by starting a helper and restarting its own unit has a trap: the helper starts inside the service’s cgroup, and systemd kills everything in that cgroup on restart. In a throwaway systemd, a setsid’d helper vanished before its last log line, KillMode=process kept it alive but leaked it into the next instance, and systemd-run gave it its own unit. This post shows the evidence, the fix, and what I did not test.
React Native's Built-In URL Does Not Resolve '../x'
A client library shared between a website and a React Native app usually fails on the boring parts of the Web platform, not on fetch. Reading React Native 0.87.1 and Expo 57 sources shows three layers of runtime, a URL class that returns ‘/b/c/d/../x’ where the web returns ‘/b/x’, and a fetch with no response stream unless Expo replaces it. A small client that checks what it needs, tested against four simulated runtimes.
A Durable Object Alarm Retries 6 Times, Then Stops
A Durable Object has one alarm, runs it at least once, and retries a failing handler with backoff up to six times. After that nothing wakes the object again. A walkthrough of a lease-expiry ledger that survives all three: a table as the source of truth, a constructor that re-arms the alarm, and an effect that tolerates being run twice. Run on local workerd, with the documented limits quoted from Cloudflare.
Durable Object Hibernation Keeps the WebSocket and Drops Every Class Field
Cloudflare’s docs say hibernation keeps WebSocket connections open and discards in-memory state. This post runs a small Durable Object under local workerd and watches exactly what that means: the socket stays OPEN, the constructor runs again on the next message, class fields reset, attachments survive only if you serialize them, and a pending timer or the standard WebSocket API quietly turns hibernation off. Production behaviour and billing are not tested.
APNs Stores One Pending Notification per App, So Treat Push as a Hint
Apple and Google both document what happens to a push when the device is offline, the app is killed or the sender is too chatty: messages are replaced, dropped, reordered or delayed. A table of those documented failure modes, and a small simulation showing that a client which pulls from a cursor converges where a client which applies push payloads does not.
Old JWT Verifiers Ignore a New Restricting Claim, So Restrict by Narrowing scope
The JWT specification tells verifiers to ignore claims they do not understand. That is harmless for claims that grant things and dangerous for claims that restrict them. Four ways to cap what a client kind may hold, tested with a small issuer and three verifiers, and the one I would pick.
A Step-Up Challenge Is a 401, and Your Client Needs a Loop Guard
RFC 9470 lets an API tell a client that its access token was obtained with too weak or too old a login. Reading the RFC section by section, then building a toy server and client, shows what the document specifies (a 401 challenge, two parameters) and the three things it leaves to you: concurrent challenges, loops, and caches.
UDP connect() Sends Nothing, and getsockname() Gives You Your Local IP
A machine with Wi-Fi, a VPN and a container bridge has several addresses, and a hostname lookup can return the wrong one. Connecting a UDP socket to a destination makes the kernel pick the source address it would use, and getsockname() reads it back without a single packet being sent. Checked on Linux with three interfaces: five destinations, five answers, no IPv4 or ARP frames on the wire.
Socket Activation Keeps Connections Waiting, Not Refused, During a systemd Restart
A runbook for replacing a running daemon without dropping requests, tested on a small TCP server under real systemd. A plain restart refused or reset connections every time; starting the new version beside the old one avoided refusals but still reset a connection in some runs, and a kernel setting made those resets go away; socket activation produced no errors. Also covered: draining in-flight work against TimeoutStopSec, and a release swap that checks the new version and rolls back.
Expo's Fingerprint Runtime Version Changes When You Edit extra, Not Your JavaScript
An over-the-air update is JavaScript that calls into a native binary you cannot change. Treat the runtime version as the ABI version of that binary. Experiments with @expo/fingerprint 0.20.13 show which edits change the hash (extra, version, build numbers, native modules) and which do not (JavaScript, pure-JS dependencies), and a skip list that silently drops a default.
systemd Kills the Updater Your Service Started, Even With setsid
A daemon that updates itself by starting a helper and restarting its own unit has a trap: the helper starts inside the service’s cgroup, and systemd kills everything in that cgroup on restart. In a throwaway systemd, a setsid’d helper vanished before its last log line, KillMode=process kept it alive but leaked it into the next instance, and systemd-run gave it its own unit. This post shows the evidence, the fix, and what I did not test.
React Native's Built-In URL Does Not Resolve '../x'
A client library shared between a website and a React Native app usually fails on the boring parts of the Web platform, not on fetch. Reading React Native 0.87.1 and Expo 57 sources shows three layers of runtime, a URL class that returns ‘/b/c/d/../x’ where the web returns ‘/b/x’, and a fetch with no response stream unless Expo replaces it. A small client that checks what it needs, tested against four simulated runtimes.
A Durable Object Alarm Retries 6 Times, Then Stops
A Durable Object has one alarm, runs it at least once, and retries a failing handler with backoff up to six times. After that nothing wakes the object again. A walkthrough of a lease-expiry ledger that survives all three: a table as the source of truth, a constructor that re-arms the alarm, and an effect that tolerates being run twice. Run on local workerd, with the documented limits quoted from Cloudflare.
Durable Object Hibernation Keeps the WebSocket and Drops Every Class Field
Cloudflare’s docs say hibernation keeps WebSocket connections open and discards in-memory state. This post runs a small Durable Object under local workerd and watches exactly what that means: the socket stays OPEN, the constructor runs again on the next message, class fields reset, attachments survive only if you serialize them, and a pending timer or the standard WebSocket API quietly turns hibernation off. Production behaviour and billing are not tested.
APNs Stores One Pending Notification per App, So Treat Push as a Hint
Apple and Google both document what happens to a push when the device is offline, the app is killed or the sender is too chatty: messages are replaced, dropped, reordered or delayed. A table of those documented failure modes, and a small simulation showing that a client which pulls from a cursor converges where a client which applies push payloads does not.
Old JWT Verifiers Ignore a New Restricting Claim, So Restrict by Narrowing scope
The JWT specification tells verifiers to ignore claims they do not understand. That is harmless for claims that grant things and dangerous for claims that restrict them. Four ways to cap what a client kind may hold, tested with a small issuer and three verifiers, and the one I would pick.
A Step-Up Challenge Is a 401, and Your Client Needs a Loop Guard
RFC 9470 lets an API tell a client that its access token was obtained with too weak or too old a login. Reading the RFC section by section, then building a toy server and client, shows what the document specifies (a 401 challenge, two parameters) and the three things it leaves to you: concurrent challenges, loops, and caches.
After Backgrounding, a WebSocket Reporting OPEN Is Only a Claim
Apple’s and Android’s own documentation say a backgrounded app can be suspended, its network access deferred, and its existing connections closed. So when your app comes back, readyState === OPEN is only a memory of the last event, not a measurement. This post derives a small foreground routine (probe, rebuild, catch up) from those documented rules, runs it against a frozen-process stand-in on Linux, and is explicit about what no device was used to check.
A WebSocket Origin Check Blocks Other Websites, Not Other Clients
The Origin header on a WebSocket handshake is set by browsers and can be set to anything by every other client. So an allow-list protects a cookie-authenticated browser session from other sites, and it proves nothing about who is connecting. A runnable server, a real headless-browser attack, and the handshake rule that follows.
Cancel the Stream, Not the Connection, When One HTTP/2 Request Times Out
A request deadline fired on a connection that carries many requests at once. Do you close the connection, or only the request? In a small Node lab, closing the HTTP/2 session failed two innocent requests, while cancelling only the stream let them finish on the same TCP connection. The same lab shows the one case where cancelling is not enough: a lost packet stalls every stream on a TCP connection, and a connection-level PING is the right way to tell.
Waiting for bufferedAmount Protects the Sender, Not a Slow Consumer
WebSocket send() never waits, and bufferedAmount only sees the part of the path that is in your own process. In a small lab, a sender that politely waited for bufferedAmount to drain still let a slow consumer queue nearly all 400 messages in its own memory, while a credit window of 8 kept the backlog at 8. This post goes through four common beliefs about send(), bufferedAmount, message size and backpressure, tests each one, and ends with a table for choosing between them.
Resuming an Event Stream with a Cursor, a Bounded Log and a Snapshot
How a client catches up after a reconnect without losing state, without duplicates, and without re-reading history. Built in steps (versions 0 to 4), from “read everything again” to a bounded log with a snapshot fallback, with what a real browser’s EventSource does on reconnect and on a non-200 response, a snapshot-ordering bug, and a decision tree.
Two write() Calls Can Stall a TCP Connection for 40 ms
Send a header and a body as two small write() calls, then wait for a reply, and every round trip can stall for about 40 ms on an idle machine with an idle network. Neither Nagle’s algorithm nor delayed ACK is a bug; together they deadlock until a timer fires. This post predicts the stall, reproduces it with a 50-line script, reads it off a packet trace, and compares the four ways out.
Enable Refresh Token Rotation and Parallel Requests Can Log Users Out
Refresh token rotation turns a reused token into a theft alarm. A client that fires three requests with an expired access token and refreshes once per 401 presents the same refresh token three times, and trips that alarm on itself. This post breaks a naive client on purpose, fixes it with single-flight refresh and a stale check, and compares the alternatives: a server grace window, request gating, proactive refresh and sender-constrained tokens.
Retries Duplicate Your Writes, and Exactly-Once Won't Save You
When a request times out, the client cannot tell whether the request or only its acknowledgement was lost. This is why exactly-once delivery cannot be built, and why effectively-once processing is at-least-once delivery plus a receiver that deduplicates. A runnable experiment with a flaky network, the bug that still duplicates 89 of 200 requests, an atomic dedupe store, and a jitter simulation.
Edge-Triggered epoll Hangs After a Partial Read, and Event Streams Break the Same Way
epoll’s edge-triggered mode stalls after a partial read; level-triggered mode cannot. The same distinction decides whether a realtime client survives lost, duplicated and reordered events. A reproducible epoll experiment, a five-way simulation on a lossy channel, and a 40-line invalidator that handles the race everyone misses: an event that arrives while the refetch is running.
A Dead TCP Peer Goes Unnoticed for 15 Minutes on Writes, Forever on Reads
TCP never tells you that the peer is gone. A reader blocks forever, a writer keeps retransmitting for a quarter of an hour, and SO_KEEPALIVE does nothing while data is unacknowledged. This post builds a one-file Linux lab with no root, climbs a ladder of mechanisms (nothing, keepalive, TCP_USER_TIMEOUT, an application heartbeat), measures how long each takes to notice, and shows where NAT and the RFCs fit in.
When a WebSocket Outlives Its Token
After the 101 response, no request carries credentials: the server remembers a verdict, not the evidence. This article follows one token through a browser handshake, shows what a page can and cannot see, and compares three ways to renew it (reconnect, in-band, out-of-band) with a runnable server, tests, and a headless-Chrome check.
Replace Polling with Change Subscriptions, and Keep the Poll as the Fallback
How a mobile client moved from timer-driven polling to shared change subscriptions over a per-message-billed relay. The subtle part is not the stream. It is deciding honestly when the fallback poll may stand down.
Batching and Bounded Frames on a Metered Relay
When every WebSocket message costs money and every frame has a hard cap, you need batching with independent flush triggers, a single-frame threshold below the cap, bounded chunking, and capability flags that survive mixed-version rollouts. The sharp edges are ordering on shutdown and a zero value that disappears from the wire.
Liveness Without Pings, and Idle Sleep
A fixed-interval keepalive costs a message in each direction even when the connection is busy. Treating every inbound frame as proof of life, probing only quiet connections, and letting an existing heartbeat tell an idle host when to disconnect removes most of that traffic, at the price of a bounded wake-up delay.
When the Close Event Never Comes
A client that waits for its own socket’s close event before cleaning up can stall silently when that event never arrives. Settle your state when you decide to end the connection, make the cleanup idempotent, and keep the event for closes you did not start.
Separate the Live Document from the Saved Copy
Designing collaborative document editing with conflict-free synchronization, saved snapshots, and safeguards against overwriting content before initial sync.
One Open Admission While a Start Is Unfinished
Controlling distributed task starts with one unfinished admission per workspace surface, a fixed destination, and idempotent retries during machine outages.
Separate Authentication from What an Actor May Do
Designing API authentication and role-based authorization with actors, route permissions, and separate credentials for people and machines.
Do Not Load the Whole Past into the Model
Designing AI agent memory with retrieved context, durable task records, and evidence-based create or update decisions that tolerate search index lag.
A Transcript Is Not a Task List
Designing a meeting AI pipeline with separate transcription, summarization, and task extraction, preserving evidence and resolving deadlines from meeting time.
Keeping an Agent Resident Behind NAT
Designing a resident agent behind NAT with outbound connections, separate liveness checks, and on-demand SSH sessions for interactive access.
Hardening a Go Daemon and a Python API
Five techniques from a polyglot codebase: rolling out golangci-lint on existing Go code, running the race detector in CI, bounding subprocess lifetimes, re-registering launchd services without the bootout race, and choosing PostgreSQL row-lock strength around foreign keys.
AI Agent Clean Architecture
Why I built an AI agent platform in Go instead of using LangChain, and how Clean Architecture made model, memory, tool, and streaming concerns independently replaceable.
Hello World
Welcome to my blog – a space for sharing notes on CS, ML, and engineering.