Skip to content

Canonical message format ​

This is the format sesh is built on: a provider-neutral, durable representation of a conversation — what was said, which tools ran, with what inputs and results — decoupled from whichever vendor system (Claude Code, Codex) produced it. It's what makes it possible to point sesh at one provider's session file and get the other provider's session/rollout file back out, with nothing lost silently along the way.

The format lives in src/messages/ in the clanker repo/library that sesh ships from, but it's useful on its own: the same thread can be referenced anywhere a conversation needs to outlive the CLI that started it — a desktop app syncing with an agent, a mobile app checking in casually, another tool picking up a task later. It's meant as the substrate a memory/context system, a storage layer, or a UI gets built on top of.

The transformation, at a glance ​

The format's whole job is to sit in the middle of a lossy-looking conversion and make it not lose anything. Every vendor shape ingests down to the same Thread/Message[], and every emitter builds a vendor shape back up from it — the arrows run both ways:

text
 Claude Code session (tree JSONL)          Codex rollout (linear JSONL)
        |            ^                            |            ^
 parseClaudeSession   emitClaudeSession    parseCodexRollout    emitCodexRollout
        |            |                            |            |
        v            |                            v            |
        +----------------->  Thread + Message[]  <-----------------+
                        (provider-neutral, v1 schema)

parse* and emit* both return a report ({ dropped, notes } / EmitReport) alongside the data — see Provider adapters for the block-by-block mapping in each direction, and Session handoff for the full round trip: a live session converted, handed to the other provider, and merged back.

The model ​

Everything lives in src/messages/format.ts. A thread is a conversation; messages belong to a thread and carry ordered blocks of content:

ts
type Thread = {
  v: 1;                                // schema version, on every stored record
  id: string;
  agentId?: string;                    // clanker agent, when applicable
  parentThreadId?: string;             // sidechain / subagent threads
  originMessageId?: string;            //   ...spawned from this message
  title?: string;
  createdAt: string;                   // ISO 8601
  metadata?: JsonObject;               // cwd, gitBranch, model, originator, ...
};

type Message = {
  v: 1;
  id: string;                          // ours; source id kept in provenance
  threadId: string;
  turnId?: string;
  parentId?: string;                   // preserved from tree-shaped sources
  role: "user" | "assistant" | "system" | "tool";
  ts: string;
  runId?: string;                      // correlation to a clanker run
  content: Block[];
  usage?: { inputTokens?: number; outputTokens?: number; cacheReadTokens?: number };
  provenance?: Provenance;
  raw?: JsonValue;                     // lossless escape hatch (opt-in)
};

type Block =
  | { type: "text"; text: string }
  | { type: "reasoning"; text?: string; redacted?: boolean }
  | { type: "tool_call"; callId: string; name: string; input: JsonValue }
  | { type: "tool_result"; callId: string; output: JsonValue; isError?: boolean }
  | { type: "file_change"; callId?: string; path?: string; patch?: string }
  | { type: "web_search"; callId?: string; query: string; results?: JsonValue }
  | { type: "attachment"; mediaType?: string; uri?: string; data?: string }
  | { type: "unknown"; sourceType: string; raw: JsonValue };   // forward-compat

type Provenance = {
  provider: "claude-code" | "codex" | "clanker" | (string & {});
  channel: "live" | "session-file";
  sessionId?: string;                  // vendor session id
  sourceId?: string;                   // vendor message id (uuid / item id)
  sidechain?: boolean;
};

A few things this shape encodes on purpose:

  • Threads span runs. Messages carry runId for correlation to a particular clanker run, but the thread itself is durable across many runs.
  • Claude is a tree, Codex is linear. Canonical storage is linear per thread (parentId is preserved when the source is tree-shaped), and sidechains / subagent transcripts become child threads (parentThreadId + originMessageId), not messages folded into the parent.
  • Turns are explicit in Codex, implicit in Claude. Adapters assign a canonical turnId in both cases so downstream consumers don't need to know the source vendor's convention.
  • Reasoning may be opaque. Codex's reasoning is usually encrypted_content-only; the reasoning block supports redacted: true with no text rather than fabricating content.

MessageStore ​

Storage is pluggable. src/messages/store.ts defines the interface plus two reference implementations (JsonlMessageStore — append-only JSONL — and MemoryMessageStore). A gateway is expected to bring its own DB-backed store later, injected the same way every other gateway concern is:

ts
interface MessageStore {
  appendMessage(msg: Message): Promise<void>;
  putThread(thread: Thread): Promise<void>;
  getThread(id: string): Promise<Thread | undefined>;
  listThreads(filter?: { agentId?: string; parentThreadId?: string }): Promise<Thread[]>;
  getMessages(threadId: string, opts?: { afterId?: string; limit?: number }): Promise<Message[]>;
}

Principles ​

  • Canonical fields for what is universal; raw + unknown for what is not. Never flatten to the lowest common denominator, and never lose data silently — an unrecognized source record becomes a { type: "unknown" } block rather than being dropped.
  • raw is opt-in per adapter ({ keepRaw: true }), since some vendor payloads (e.g. Claude's rich toolUseResult) can be large and most callers don't need the original record kept around.
  • Adapters and emitters return reports, never fail silently. Ingesting a session reports what was skipped (telemetry, session-state lines); emitting a session reports what was dropped or transformed. See Provider adapters and Session handoff.

What's not in v1 ​

  • Storage backends beyond the JSONL reference implementation and the in-memory one — a gateway injects its own.
  • Cross-thread search, embeddings, or memory — that's a layer built on top of this format, not part of it.