This video explains how AI usage limits are consumed in ChatGPT and Claude, focusing on the role of caching. It provides practical tips to reduce token usage, such as avoiding long sessions, using Claude Projects for caching, and understanding cache lifetime and model changes.
A short editorial from the VEONIB team on why this content matters.
The video effectively demystifies AI caching, showing how simple habits like session management and project use can drastically cut token consumption.
Unlike generic tips, it dives into technical nuances like cache lifetime and model-specific behaviors, offering actionable insights for power users.
AI power users and professionals should watch this to optimize their ChatGPT and Claude usage; then implement caching strategies to extend limits.
Restrictions on how much you can interact with AI models like ChatGPT or Claude within a subscription plan.
A mechanism where AI models reuse computations for identical input prefixes to reduce processing costs.
Strategies to minimize token consumption in AI interactions to extend usage limits.
A feature in Claude that allows uploading documents to a project context, which are cached for efficient reuse.
A mode in ChatGPT designed for longer tasks, with separate usage limits from the standard Chat.
The duration for which cached computations are stored; if exceeded, cache is cleared and must be recomputed.
A parameter in some AI models that controls how much effort the model puts into a task, affecting token usage and cache behavior.
Secondary AI agents used within a session that maintain separate caches, potentially increasing overall token usage.
Why does my ChatGPT or Claude usage limit run out so quickly?
Because each message in a session resends the entire conversation history, consuming tokens for all previous messages.
How can I reduce token usage in long conversations?
Start a new session when previous context is not needed, and batch multiple questions into one input.
What is caching in AI models?
Caching stores computations for identical input prefixes so that subsequent requests with the same prefix can reuse them, saving tokens.
How does Claude Projects help save usage limits?
Uploaded documents in Projects are cached; when reused, they don't count toward usage limits, reducing token consumption.
Why does changing the model mid-session break caching?
Because the cache is tied to the specific model's computations; a different model requires recomputation, invalidating the cache.
What is the cache lifetime for Claude and ChatGPT?
Claude Code caches for up to 1 hour; ChatGPT Work and Codex likely have a minimum of 30 minutes based on model specs.
Does using the compact command save tokens?
No, frequent use of compact can break caching and may increase token usage because it reprocesses the session.
Why do subagents increase token usage?
Subagents maintain separate caches, so using them can lead to more overall token consumption than a single agent with caching.
How do external tools like MCP affect token usage?
They add extra AI processing steps (tool search and execution), increasing the number of inputs and thus token usage.
Should I change the effort level mid-session?
It depends on the model; for some (e.g., Claude 5.1), it's safe, but for others (e.g., GPT-5.6), it may break caching and increase costs.