Read the stream,
not the answer.

every-other-token is a Rust CLI and web UI that sits on an OpenAI or Anthropic token stream as it arrives. It rewrites every other token (or any fraction), and it tells you how sure the model was about each one: confidence exp(logprob) and perplexity exp(-logprob), per token, live.

transform

what the model sent

what you see

0 tokens 0 rewritten mean confidence - sure ≥70%unsureguessing <40%

Recorded from the real /stream endpoint of every-other-token --web --provider mock. The mock provider replays a fixed reply with fixed logprobs through the same interception pipeline a real model goes through, so it runs with no API key. Point it at OpenAI or Anthropic and the tokens and numbers are the model's own.

00How it works

Four stages, shown on a real run of the mock provider with the reverse transform at rate 0.5.

Diagram: 1 Intercept reads the live stream, 2 Score gives each token confidence exp(logprob) and perplexity exp(-logprob), 3 Mutate reverses the odd-numbered tokens, 4 Output shows 'The kciuq brown xof jumps revo the yzal dog' in the terminal, web UI or JSON
  1. Intercept. It opens the model's streaming (SSE) connection itself, so it sees each chunk the moment it arrives.
  2. Score. Each token gets confidence exp(logprob), perplexity exp(-logprob) and the top alternatives.
  3. Mutate. Every other token by default (--rate), or only the ones below --min-confidence, goes through a transform.
  4. Output. Terminal, web UI, JSON lines, CSV or an HTML heatmap.

01Running in two minutes

No Rust needed: Windows .exe (double-click it and the web UI opens), macOS Apple Silicon, macOS Intel, Linux x86_64. Or scoop install every-other-token after scoop bucket add mattbusel https://gitlab.com/mattbusel/scoop-bucket, brew install mattbusel/tap/every-other-token, or cargo install every-other-token (Rust 1.81+).

install
$ cargo install every-other-token
try it, no API key
$ every-other-token "What is consciousness?" --provider mock --visual
a real model, in the browser
$ export OPENAI_API_KEY=sk-...   # or ANTHROPIC_API_KEY with --provider anthropic
$ every-other-token --web        # http://localhost:8888

02In the terminal

The stream prints as it arrives. With --visual, rewritten tokens are highlighted; --heatmap colors each token by importance instead.

Terminal: every-other-token with the mock provider, every other word reversed and highlighted, 22 tokens streamed, 11 transformed
Real output of the command above.

03Every token carries its numbers

Each intercepted token is an event with the original text, what was shown, its index, whether it was rewritten, confidence, perplexity and the top alternatives. Consume them from the library, as JSON lines with --json-stream, or export them to CSV, JSONL or a self-contained HTML heatmap.

Per-token table from cargo run --example mock_stream: index, original token, shown token, confidence, perplexity
cargo run --example mock_stream, from a clone of the repo.

04The web UI

--web serves a single page with no build step and no external scripts. Single, split, quad (four transforms at once), OpenAI vs Anthropic diff, A/B system prompts, a research dashboard, JSON and CSV export, and collaborative rooms where several people edit tokens mid-stream.

The web UI in split view: original stream on the left, transformed on the right, each token underlined by confidence, with perplexity and confidence sparklines below
Split view, mock provider. The underline under each token is its confidence.

05What else is in the box

9 transforms, chainable

reverse, uppercase, mock, noise, chaos, scramble, delete, synonym, delay:N, or a chain like reverse,uppercase.

Rate, seed, gating

--rate 0.3 spreads rewrites evenly at any fraction. --seed makes random transforms reproducible. --min-confidence only touches tokens the model was unsure of.

Provider diff

--diff-terminal streams OpenAI and Anthropic side by side and compares their confidence structure.

A/B system prompts

Run two system prompts against the same user prompt, many times, and test the confidence shift with Welch's t-test (--significance).

Research mode

--research --runs 20 runs headless and writes aggregate perplexity, confidence and vocabulary stats to JSON.

Record and replay

Record any session to JSON and replay it later, for example to spot behavior changes after a model update.