Will it fit, and how fast?

Enter a document size in any unit. Compare stated context capacity, a conservative working zone, and illustrative response time across the included models.

10 models Verified 2026-05-06 Latency is illustrative - run your own evals

Runs entirely in your browser. No signup.

Stated context windows are technical ceilings. Real reliability degrades past ~50% of the window (Chroma “Context Rot” research). The reliable working zone is shaded.

Sources & methodology.

Every number this tool shows, and where it comes from.

Context windows. Sourced from vendor documentation: Anthropic, OpenAI, Google AI for Developers, Together AI, DeepSeek API.

Reliable working zone. 50% of the stated window, based on Chroma Research’s “Context Rot” findings (2025) and the broader long-context-degradation literature. Some models hold reliability further; we err conservative.

Token conversion. ~4 chars/token, ~0.75 words/token, ~500 tokens per page of typical English prose, ~250 tokens per KB. These are averages; technical writing and code can run 30-40% denser.

Latency assumptions. TTFT and tokens-per-second are illustrative defaults derived from public benchmarks and our own measurements. Real-world latency depends on region, traffic, and prompt size. Always validate against your own region under your own load.

Cache hit reduction. Warm-cache TTFT is approximated as ~70% of cold TTFT. Anthropic and OpenAI both ship faster TTFT on cache hits, with the exact ratio depending on prompt structure.

Model set last verified 2026-05-06.

Keep measuring.

The rest of the toolkit runs in your browser too.

Choose with the workload in view.

Open Orbit to keep model choice, cost, approvals, and run history together.

500 trial credits. Card required. $0 charged today.