\n\n\n\n Nine Ways to Await and Every One of Them Bites - AgntHQ \n

Nine Ways to Await and Every One of Them Bites

📖 5 min read•823 words•Updated Sep 12, 2026

It’s 11:40 on a Tuesday night. You asked your coding agent to add a background log write to a service handler, it produced twelve lines that look exactly like the async code you’ve read a thousand times, and the tests pass. You ship it. Three days later, a chunk of your logs are missing, and nobody can reproduce it locally. The function ran. The first half of it, anyway.

If you’ve lived that, you already have an intuition for what Gavin Gray, Shriram Krishnamurthi, and Will Crichton put on paper in A Design Space Exploration of Async/Await, out of Brown’s Cognitive Engineering Lab and headed to OOPSLA 2026. Their claim, stripped of academic politeness: async and await are not one feature. They’re a family of features wearing the same two keywords, and the differences between family members are big enough to change whether your code runs at all.

The pitch versus the reality

The selling point of async/await has always been straight-line asynchrony. You write code that reads top to bottom, the runtime handles the waiting, and you skip callback nesting and manual state machines. That’s a real win, and it’s why the syntax spread across so many languages so fast.

The paper’s contribution is mapping out nine design dimensions where languages diverge underneath that shared surface. Two of those dimensions matter more than the rest for anyone writing production code: task lifecycle and cancellation. When does an async function actually start executing? At the call, or only when you await it? What happens to a task nobody is awaiting? Can it be killed mid-flight, and if so, at which points, and what runs on the way out?

The paper’s own illustration is almost insultingly simple. An async function prints “A”, awaits a simulated two-second log write, then prints “B”. Whether “B” ever prints depends on decisions the language designer made years before you typed the function. Same keywords, same shape, different outcome.

Why this is an AI tooling problem, not just a language nerd problem

This is where I get to be the killjoy. Coding agents are pattern machines trained on an enormous pile of code from every language that shipped these keywords. The syntax is nearly identical across all of them. The semantics are not. That’s the worst possible setup for a statistical model: maximum surface similarity, maximum hidden divergence.

So when your agent writes async code, it’s drawing on a blend of conventions where the shape is shared and the meaning isn’t. Most of the time you get something that compiles and mostly works, because the happy path across these languages does converge. The trouble shows up in exactly the places the paper flags: a task that’s constructed but never started, a task dropped at a cancellation point you didn’t know existed, cleanup that doesn’t run because the language doesn’t guarantee it will.

Those bugs share an ugly profile. They don’t fail loudly. They fail as missing writes, occasional timeouts, resources held slightly too long, behavior that changes under load. The kind of thing you don’t catch in review because the code looks correct, and the kind of thing that survives a test suite because your tests don’t exercise the cancellation path.

What to actually do about it

I’m not going to tell you to stop using agents for async code. That’s not useful advice and you weren’t going to take it. Try this instead:

  • Treat async code from an agent as the highest-scrutiny category of output you accept. Above SQL. Above config. Read it line by line.
  • Ask the specific questions: does this task start eagerly or lazily in this language? What happens if the caller goes away? Does cleanup run on cancellation, and is that a guarantee or a convention?
  • Write a test for the cancelled path, not just the completed one. It’s the case nobody writes and the case that breaks.
  • Be suspicious when your agent moves between languages in one session. The keywords carry over. The behavior does not.

Note that the paper landed on Hacker News and picked up nine points. Nine. A research team spent real effort articulating a design space that explains a whole category of bugs sitting in production right now, and it barely registered against whatever model release ate the front page that day. That ratio tells you something about where developer attention goes and where it probably should go instead.

My take

Async/await did simplify concurrent programming, and I’d rather write it than manage callbacks by hand. But “simpler to write” and “simpler to reason about” came apart somewhere, and the gap widened once code generation entered the picture. A tool that produces plausible async code faster than you can verify it is not saving you time; it’s moving the cost to a worse place, at a worse hour, with less context.

Nine dimensions. Two keywords. Read the paper, then go read your own async code with fresh suspicion.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top