> Warden / blog
• workspace, control, usability

A long session that survives a hard stop

Warden was run through hundreds of live tool rounds, killed mid-task, and resumed with its session intact. The rough edges the same verification found are fixed.

Long coding sessions rarely fail in dramatic ways. A process gets killed. The context fills up and gets summarized at an awkward moment. A request is cancelled, and nobody is quite sure it stopped. A session that cannot ride out those moments is not one you can leave running, so Warden’s pre-release verification spent real time there.

Killed mid-task, then resumed

Warden ran 279 tool rounds against a live model across three long runs: a read-only investigation of its own codebase, a write-heavy run in a scratch project that fixed 12 failing tests, and a run with a deliberately small context window that had to compact five times. Along the way the process was killed outright, once between turns and once mid-turn, right after a compaction. Both times, resuming brought back the same session with its full transcript, and the model knew what it had been doing. Memory use leveled off at about 300 MB and stayed there.

When handing off a paused session came up here before, resume was framed as continuity rather than perfect recreation. That framing stands. What has changed is that a resumed session now keeps its identity and adds to its own history.

Rough edges, found and fixed

The same verification turned up rough edges, and those are fixed. One came straight out of the long runs: after an automatic compaction in the middle of a task, the model could treat the summary as the end of the job. The summary now makes clear that the work continues, so the model carries on. A cancelled request is never quietly sent again, so stopping a turn does not keep spending. Rewinding a turn deletes files the turn created and returns files its edit and write tools changed to their pre-turn content, even a file edited twice in one turn. A resumed session also takes its reading of how full the context is from the provider’s own count, not from a fresh estimate.

None of this is glamorous. It is what makes it reasonable to leave a long task running and come back to it later.

Warden is a working proprietary development build. Supported platforms and session behavior will be stated with its public release.

Was this useful?