THE SEQUENTIAL BOTTLENECK
Here's how most people use coding agents: give it a task, wait, give it the next task. Sequential by default. Not because the tasks depend on each other, but because the tooling assumes one thread.
We started noticing the pattern in our own work. Refactoring three modules. Running tests while fixing a bug elsewhere. Researching an API while scaffolding the integration. Independent tasks, sitting in a queue, waiting their turn for no reason.
"Just open more terminals" is the naive fix, and it's what we did at first. It creates a different problem. Each session is blind to the others. No shared state, no aggregate cost tracking, no way to see what six agents are doing at a glance. You become the orchestrator, mentally tracking which tab is doing what, copy-pasting context between sessions.
So we built the orchestrator.
ARCHITECTURE
Flock is a native macOS application built in Swift and AppKit. It manages multiple Claude Code CLI instances as isolated subprocesses, each with its own PTY, environment, and context window.
PROCESS ISOLATION
Each agent session spawns a separate claude process with --output-format stream-json --verbose. The process gets its own:
- Pseudo-terminal (PTY) via SwiftTerm
- Shell environment with isolated ZDOTDIR (prevents cross-contamination)
- Bidirectional I/O pipes for input injection and output parsing
- Independent working directory
Process isolation is critical. If two agents are editing files in the same repo, their tool calls must not interfere with each other's context windows. Each process maintains its own conversation state. The multiplexer orchestrates, but never merges contexts.
STREAM PARSING
The stream-json output format gives us structured events for every action the agent takes: thinking blocks, tool calls, tool results, text responses. We parse these in real-time on a dedicated queue.
stdout → readabilityHandler → parser queue → event dispatch
↓
AgentTaskAction {
type: think | read | edit | write | bash | search | agent | web | message
timestamp: Date
content: String
metadata: [String: Any]
}
This gives us a structured action timeline for every running agent. We know what each agent is doing at any moment -- not from terminal output scraping, but from the actual event stream.
STATE DETECTION
Terminal output is noisy. An agent that's "thinking" looks identical to an agent that's "waiting" if you're just watching the terminal. We built a state detector that classifies agent activity in real-time:
- Active: byte rate > 150 bytes/sec (agent is producing output)
- Thinking: Claude is reasoning (detected from stream events)
- Running: a tool call is executing (Bash, Read, etc.)
- Waiting: idle for > 2.5 seconds
- Error: process exited with non-zero status
Each pane's tab shows its current state with a visual indicator. Across 6 running agents, you can see at a glance which are active, which are blocked, which are done.
AGENT MODE
Beyond raw terminal multiplexing, Flock has a dedicated agent mode -- a kanban-style interface for orchestrating multi-step workflows.
The model is simple. You create tasks. Tasks have a state machine:
backlog → inProgress → done | failed
A scheduler manages capacity. You configure maxParallelAgents (default: 3). When a running task completes, the scheduler pulls the next task from the backlog. Each task gets its own Claude process, its own action timeline, its own cost accumulator.
RESUMPTION
Claude Code sessions have a sessionId. When a task completes but needs follow-up, Flock resumes it with claude --resume <sessionId> -p <message>. The resumed process inherits the full conversation history. Cost accumulates across resumption rounds:
task.costUsd += resumedSession.cost
This means a single task can span multiple Claude sessions while maintaining continuous cost tracking and a unified action timeline.
OBSERVABILITY OVER CONTROL
The interface isn't a terminal emulator. The interesting design decision was treating agent output as a structured event stream rather than text to scroll through. You see what the agent read, what it edited, what it was thinking -- as discrete actions with metadata, not raw stdout. This matters when you're watching three agents simultaneously. You need to pattern-match on behavior, not read every line.
SHARED MEMORY
Parallel agents with isolated contexts create a coordination problem: how does agent B know what agent A discovered?
Flock's memory system bridges this gap. A MemoryStore maintains a shared pool of entries (max 200, auto-trimmed, pinned entries preserved). When memory is enabled, Flock generates a .flock-context.md file in the working directory that all agent sessions read as persistent context.
Memory entries are categorized:
- project -- facts about the codebase or current work
- preference -- how the user wants things done
- correction -- mistakes to avoid
- taskSummary -- auto-generated when tasks complete
When a task completes, its summary is automatically written to the memory store. Subsequent tasks see that summary in their context file. This creates an implicit coordination layer -- agent B doesn't talk to agent A, but it reads what agent A concluded.
TOKEN ECONOMICS OF PARALLELISM
The obvious argument for parallelism is wall-clock time. Run five tasks at once, finish in the time of the slowest one. But there's a less obvious argument that matters more: context isolation saves tokens.
Sequential workflows accumulate context. Each new task in the same conversation inherits everything from previous tasks. By task 5, you're paying for 4 tasks worth of context you don't need. Parallel agents start clean. Each window contains only what its task requires.
sequential (5 tasks, same context):
task 1: ~40K tokens
task 2: ~80K tokens (carries task 1)
task 3: ~120K tokens (carries 1+2)
task 4: ~160K tokens (carries 1+2+3)
task 5: ~200K tokens (carries 1+2+3+4)
total: 600K tokens
parallel (5 tasks, isolated):
task 1: ~40K tokens
task 2: ~40K tokens
task 3: ~40K tokens
task 4: ~40K tokens
task 5: ~40K tokens
total: 200K tokens
3x token efficiency, plus the tasks complete in wall-clock time of the slowest one instead of the sum of all five. The shared memory file adds maybe 2K tokens per context. The savings compound with task count.
USAGE TRACKING
Running multiple agents simultaneously means cost can spike fast. Flock tracks everything in real-time:
- Per-session: input tokens, output tokens, cache read/creation tokens
- Per-model pricing: Opus, Sonnet, Haiku rates applied automatically
- Plan limits: OAuth API polling for 5-hour and 7-day utilization
- Aggregate daily: total cost, session count, token breakdown
The status bar shows utilization percentage and reset countdown. You know exactly how much runway you have before hitting plan limits -- critical when you're running 3 agents simultaneously burning through your allocation.
WHAT WE LEARNED
After a few months of daily use, the takeaway isn't about the architecture. It's that most agent work is embarrassingly parallel and we just didn't know it. The tooling assumed sequential. We assumed sequential. Once parallel is as easy as opening a new pane, you realize the queue was artificial.
The other thing we didn't expect: isolation makes agents better. Agents that can't see each other's context make fewer compounding errors. No confusion from tool output belonging to a different task. No stale assumptions from three tasks ago bleeding into the current one. The shared memory file gives just enough coordination. Everything else is noise you're better off without.
Flock is open source at github.com/Divagation/flock.