uam — what's next

Four ways uam could use Copilot SDK capabilities it doesn't touch yet, each shown as it works today and as it would after.

Proposals as of .

The ideas came from a council of Claude Opus 5.5 and Claude Fable 5.1. Five brainstorm branches ran in parallel, each isolated under its own frame (on-call, speedrunner, game design, remove-the-assumption, markets). Every idea was then scored, the survivors were clustered, and each SDK call was checked against Copilot SDK v1.0.17.

Tool workbench

Re-run any tool call with edited arguments, or stop the one stuck command without aborting the turn.

Now

The agent runs npm test and it hangs on a watcher. The row spins and the turn never ends. You can abort the whole turn and lose the agent's plan, or open Background tasks, guess which shell it is and stop it; the agent then sees a bare failure and often runs the same thing again. Re-running a grep with another pattern means typing a prompt, paying for a model round trip, and hoping the agent runs it as written.

After

Right-click the stuck npm test row and choose Stop this command. The confirmation opens beside the row; type “watch mode hangs, use --ci”. The row turns Stopped by you, and the agent reads the failure plus your note and re-runs with --ci in the same turn. Right-click an earlier grep, choose Run with changes…, edit the pattern, and the result shows beside the row in under a second with no model call. Send result to agent puts it into the conversation.

How it works

The row menu gains Run with changes… and Stop this command. The reader (about 420px) sits beside the lifted row; Stop uses the anchored confirmation with an optional “Tell the agent why”. On a phone, long-press a row; the reader becomes a bottom sheet with a sticky 44px Run or Stop button above the keyboard.

  • session.tools.execute
  • tools.getBuiltinDescriptors
  • session.shell.executeUserRequested
  • session.shell.cancelUserRequested
  • session.tasks.list
  • tasks.cancel
  • OnPostToolUseFailure
  • tool.user_requested

Honest limits

  • shell.kill only kills shells started through shell.exec, not the agent's own.
  • Shell info carries no tool call id, so matching a row to its shell relies on the command text.
  • Descriptors don't cover MCP or host tools; those get a raw JSON editor.

Risk. Whether cancelling a synchronous agent shell mid-turn comes back as a normal tool failure while the turn carries on.

First step

A throwaway probe on a real CLI (gpt-6-luna, scratch folder): sleep 600, then tasks.cancel; log the complete event, the hook and the turn. Try tools.execute and executeUserRequested both idle and mid-turn.

Spin-offs

  • Stop with a reason
  • Edit, then answer a pending shell approval
  • Diff two runs
  • Move to background instead of kill
  • Pin a re-run as a per-Project quick check

Disk ledger

See which Copilot session folders belong to which Task, and reclaim the ones nothing uses.

Now

Deleting a Task says “The provider conversation on the host is untouched”, and its session-state folder stays, invisible in uam. On the test install that adds up to 343 folders and 98 MB, 321 of them orphans from tests, forks and deleted Tasks. When the disk fills, the only tools are du and rm on UUID folders, with live Tasks' folders mixed in and nothing marking which are safe.

After

Each Task shows “On disk 2.1 MB” beside its context ring; the reader splits that into transcript, checkpoints and files. Settings → Storage reads “Copilot sessions: 98 MB in 343 folders, 321 without a Task”, biggest first, with active and in-use rows locked. Tick Orphans, press Free 61 MB from 312 folders, confirm beside the button, and the rows fade out: “Freed 59.8 MB. 4 folders are still on disk.” Deleting a Task offers “Also remove its 2.1 MB Copilot folder”, off by default.

How it works

A chip in the Task head opens a reader with Transcript, Checkpoints and Files sections. Settings → Storage has filters (Orphans, Archived, Settled, All), a size-sorted table with checkboxes, a sticky foot button and an anchored confirm. On a phone: one row per folder (name or No task, size, age) and the usual phone confirmation.

  • sessions.getSizes
  • sessions.list
  • sessions.pruneOld DryRun, ExcludeSessionIDs, IncludeNamed
  • sessions.bulkDelete
  • workspaces.listCheckpoints

Honest limits

  • No per-id failure reasons from bulk delete.
  • The “cheapest to lose” ranking is uam's own.
  • Checkpoints of settled Tasks need a direct disk read.
  • uam's own disk use (attachments, charts, fork worktrees) is invisible to the SDK; uam has to measure it.
  • On the test install the gain is hygiene (98 MB), not gigabytes, and the UI should say so.

Risk. A wrong orphan classification deletes a transcript with no undo. The Task-to-session join (provider_session_id, forks, CheckInUse) must be proven before any delete button exists.

First step

A read-only GET /api/storage that joins GetSizes, ListSessions and sessions.json, checked to report 24 Tasks, 321 orphans and 98 MB on the test install.

Spin-offs

  • Delete Task takes its folder along
  • An orphan dot in Settings
  • uam's own footprint: attachments and worktrees
  • A weekly dry-run Routine with a push
  • Export before delete

Context surgery

Show what fills the context window, compact it with a focus you choose, and say when the prompt cache broke.

Now

The ring turns amber at 80%. Context usage shows buckets, but Messages is one number: nothing says the window is three bash outputs and one 40k-token file read. The only move is typing /compact and waiting; afterwards a line says “Compacted the conversation · freed 61,230 tokens”, with no summary to check and no way to steer what it keeps. When the next turn suddenly costs more, nothing says the prompt cache was lost or why.

After

The same reader lists Heaviest messages: tool: bash 18,420 (12%), tool: view 14,100 (9%), assistant 9,870… so two stale tool outputs visibly own a fifth of the window. Below: Compact, frees up to 84,100 tokens, a focus line and Compact now. The section then reads “Freed 61,230 tokens · 42 messages removed · 131,900 → 70,670 tokens occupied”, with the summary one tap away. In the transcript, “Prompt cache lost” explains the expensive turn that followed.

How it works

The existing Context usage reader (440px, opened from the ring) gains Heaviest messages (up to 10 rows, each with a 2px share bar) and Compact. The result replaces the section in place. No new spinner: the status glyph already turns during compaction. On a phone: the existing bottom sheet, two-line rows, a full-width Compact now.

  • metadata.getContextHeaviestMessages
  • session.history.compact CustomInstructions, Trigger, TokenLimit
  • prompt_cache_break event

Honest limits

  • No per-message trimming: snapshot ids aren't event ids, and truncation is Rewind, which already ships.
  • clearContext only works from inside a tool handler.
  • There's no dry-run compact, so the before number is an “up to” estimate.

Risk. The list reads as if each row can be cut, when only whole-conversation compact exists. Compacting during a busy turn is undocumented.

First step

Extend contextSession in internal/adapter/copilot/context.go with HeaviestMessages and Compact, add ?heaviest=true and POST /context/compact, then prove TokensRemoved > 0 and a SummaryContent on an idle Task.

Spin-offs

  • Compact on model switch, sized to the new model's window
  • A cache-lost tally in usage
  • An agent-side uam_clear_context host tool (needs a decision)

The non-obvious pick

Manager loadout

The main agent reads and plans; subagents on your cheaper models do the edits.

Now

A Task on a strong, expensive model reads code, edits files, runs tests and retries them itself, every turn at the top model's price. The subagent model allowlist only applies when the main model chooses to delegate. To stop the big model typing edits you write “don't edit, delegate” into each prompt and hope; in Safe mode you also approve each direct edit.

After

Tap the loadout chip and choose Manager; the transcript reads “Loadout: Manager from here”. The main model reads, plans and writes todos. For changes it starts subagents on your allowed cheaper models, tagged worker in the subagent column, which edit and run tests. The main model has no edit or shell tools, so it can't quietly do the work itself. Usage splits main-agent spend from the workers'. Switching back applies from the next turn; transcript, todos and queue are kept.

How it works

A loadout chip beside the Agent chip, disabled during a turn or while a custom agent is selected. A transcript divider marks each change. On a phone the chip (icon plus F, R or M) opens a bottom sheet.

  • DefaultAgent.ExcludedTools main agent only
  • ExcludedTools session-wide, for Read-only
  • session.options.update
  • ExcludedBuiltinAgents
  • tools.getBuiltinDescriptors
  • usage.getMetrics AgentMetrics

Honest limits

  • Manager can't switch live: options.update has no defaultAgent, so it needs a resume between turns.
  • A preToolUse deny can't stand in: the hook can't tell the main agent from a subagent.
  • Manager can't combine with a custom agent.

Risk. Whether built-in all-tools subagents (task, general-purpose) keep the edit and shell tools that are hidden from the default agent. If not, register a uam-worker custom agent with an explicit tool list. uam_create_task and sql also need closing off.

First step

A Go probe on CLI 1.0.94 with gpt-6-luna or deepseek flash: with DefaultAgent.ExcludedTools the main agent can't edit, delegation through task can, and resuming without it restores the tools.

Spin-offs

  • Live Read-only
  • Choose which built-in workers exist
  • A main-versus-workers cost line
  • A per-Project default loadout
  • The worker's brief on its row

Fixes in flight

Two smaller fixes to the subagent column, being worked on now.

Subagent rows say what each one does

Now

A predefined agent such as the built-in task agent shows the same static blurb on every spawn, “Execute development commands like tests, builds, linters, and formatters…”, so three rows read identically.

After

Each row shows what its call asked, and while it runs, its live step.

The subagent popover stops jumping

Now

The hover peek and the anchored transcript panel move while subagents run: rows re-rank as each one finishes, the peek flips between top and bottom as its steps grow, and the conversation scrolls under the open panel.

After

Built, on the dev instance for testing: while a peek or transcript is open, the rows keep their order and the conversation keeps the open row in place. The panel stays where it opened, and the peek shifts instead of flipping. Measured: peek 0 px of movement, panel 0 moves.

The council

Five isolated brainstorm branches, six ideas each. Opus 5.5 took the on-call and speedrunner frames; Fable 5.1 took game design, remove-the-assumption and markets. Every idea was scored for novelty (N), viability (V) and fit (F), weighted .35N + .40V + .25F. A trap is an idea that sounds good and fails on contact; each one carries its reason.

Context surgery plays

Convergent: 4 of 5 branches. Became Context surgery

  • Evict the fat one: heaviest message → cutnovelty 7, viability 6, fit 8on-call
  • Drop inventory: discard heavy messages, tokens freed as lootnovelty 7, viability 6, fit 8game design
  • Trimmable context ringnovelty 7, viability 6, fit 8remove-the-assumption
  • Context short-selling laddernovelty 7, viability 6, fit 8markets
  • Route timer: mark where the prompt cache brokenovelty 8, viability 7, fit 6speedrunner
  • Cache-break spot price before sendnovelty 9, viability 5, fit 6Trap: predicts something unknowable before send.

Direct tool hands

Became Tool workbench

  • Shoot the shell, spare the agentnovelty 7, viability 8, fit 8on-call
  • Tool replay without the modelnovelty 8, viability 7, fit 8speedrunner
  • Frame-perfect veto via streamed tool argsnovelty 8, viability 5, fit 7Trap: steering lands after the tool runs; Safe-mode permission already gates it.

Shape the agent: loadouts and hooks

Became Manager loadout

  • Hands-tied manager modenovelty 8, viability 7, fit 7speedrunner
  • Loadout screen: tools for main vs subagents, one-shot skillsnovelty 7, viability 6, fit 7game design
  • Per-turn tool loadout chipnovelty 8, viability 5, fit 8remove-the-assumption
  • Cheat codes: short inputs expand into promptsnovelty 6, viability 8, fit 6Trap: duplicates / commands and skills.
  • Objective-driven Tasks: edit a living goal filenovelty 8, viability 6, fit 7overlaps the native goal items shipped in #481

Continuity and handoff

Convergent: 3 branches.

  • New Game+: hand a finished Task to a fresh one on another modelnovelty 7, viability 7, fit 8game design
  • Baton Tasks: handoff + suspend + ephemeral skillnovelty 8, viability 6, fit 7remove-the-assumption
  • Handoff as a futures contract to a cheaper tiernovelty 7, viability 6, fit 7markets
  • Freeze till morning: suspend + pause the queue from a pushnovelty 6, viability 7, fit 7on-call

Guardrails, forensics and spend

  • Black-box recorder: redacted debug bundle on errornovelty 6, viability 8, fit 7on-call
  • Amnesty reset: drop tonight's approvalsnovelty 6, viability 8, fit 6on-call
  • Standing orders: persisted “always allow” rules per Projectnovelty 6, viability 8, fit 8markets
  • Downshift instead of kill on predicted credit burnnovelty 7, viability 5, fit 8Trap-ish: a silent model switch changes behaviour.
  • Mana bar: turns left before the credit limitnovelty 7, viability 6, fit 8game design
  • Death warp: auto-rewind and retry on another modelnovelty 7, viability 5, fit 6Trap: silently discards work.
  • Combo meter for tool streaksnovelty 6, viability 8, fit 3Trap: noise, and it breaks the one-ring rule.

Housekeeping and typed outcomes

Disk rent became Disk ledger

  • Disk rent and liquidationnovelty 7, viability 8, fit 8markets
  • Typed finish line: JSON commit message / PR title → Git panelnovelty 7, viability 7, fit 8speedrunner
  • Checkpoint scrubber timelinenovelty 7, viability 7, fit 8remove-the-assumption
  • Fleet-wide steering order booknovelty 7, viability 5, fit 6Trap: fights uam's own prompt queue, which stays.
  • Mid-turn steering through the CLI queuenovelty 5, viability 7, fit 8Trap: the same conflict.

What if a Task could hand its baton to a fresh Task on a cheaper model, with history.summarizeForHandoff and session.suspend, so “long Task” stops being a thing?

The provocation the council left open.