- Lab
-
Libraries: If you want this lab, consider one of these libraries.
- AI
Token Management and Cost Optimization with Claude Code
In this lab, you will learn to manage Claude Code’s token usage, session context, and model selection to balance cost, performance, and output quality during agentic coding workflows.
Lab Info
Table of Contents
-
Challenge
Introduction
In this lab, you’ll explore how token usage and session context affect the cost, speed, and quality of Claude Code workflows. Through hands-on activities, you’ll practice monitoring usage, managing context, selecting appropriate models and effort levels, and balancing token efficiency with task quality.
Claude has already been set up, and some code and configuration files were provided, to help guide you through the foundations of token management and cost optimization.
-
Challenge
Token and session overview
Tokens and session context
A token is a unit of text processed by a language model. A token may be a word, part of a word, punctuation, or whitespace. Token counts therefore do not map one-to-one to word counts.
For each turn, Claude Code builds an active context that can include the following:
- system and project instructions;
- the conversation history that has not been cleared or summarized;
- file contents and tool results brought into the session;
- your current prompt; and
- space for Claude's response and reasoning.
Conceptually, this active context accompanies each new request. Prompt caching can make repeated prefixes cheaper and faster, and Claude Code may clear old tool results or compact history, but retained context still occupies the working window. This is why a short new question late in a long session can consume far more input tokens than the same question in a fresh session.
One way to understand this at a high-level is as follows:
turn cost ≈ retained input context + new input + generated output + reasoning
The exact bill depends on the model, authentication plan, caching, provider, and organization pricing. 1. At the Claude Code prompt, enter
/context-
Then examine the usage with
/usage -
Press
escto close the view. -
Then prompt Claude with the following:
Read README.md and explain the project in two sentences. Do not inspect other files. -
Run
/contextand/usageagain, and note the change in session metrics.
/usageshows session token details and a locally computed cost estimate for API usage. Subscription users may see plan-usage information instead; an estimated dollar value is not necessarily their billed amount. ### Growing context mattersAs active context grows:
- Cost can rise: more input must be processed on subsequent turns, although cached tokens may be priced differently.
- Responses can slow: larger prompts take more time to process.
- Quality can decline: stale or conflicting details can distract the model, and the most relevant evidence can become harder to prioritize.
- Compaction approaches: Claude Code warns as the auto-compact threshold nears and eventually summarizes older history to free space.
Warning signs include slower turns,
/contextshowing little free space, an auto-compact warning, repeated confusion with old requirements, or rising input totals for otherwise similar tasks. A context warning is not the same as a subscription or spending limit. -
Challenge
Minimizing processing tokens
Specific prompts reduce irrelevant exploration and unnecessary file content.
-
Run
/usage, and grab a baseline for the comparisons you will do next. -
Compare these approaches:
Read the entire repository and tell me how tax works.andRead src/pricing.py and tests/test_pricing.py. Explain how tax is calculated. -
Using
/usageto see how rapidly the token limit raises from a broad search versus a narrowed seach
Note: broad requests can still be benefitical in cases of architecture discovery, but should be used sparingly.
-
-
Challenge
Preserve recurring context with CLAUDE.md
A project-root
CLAUDE.mdcontains recurring project instructions that Claude Code loads at the start of sessions. It avoids repeatedly typing conventions and can survive compaction because the root file is re-read. It still consumes context, so keep it concise and reserve it for durable information rather than temporary task notes.-
To the right of the Terminal, click the +, select New tab, then click Terminal.
-
In the new Terminal, at the
workspace $prompt, runcd claude_token_lab.Then run the following:
cat > CLAUDE.md <<'EOF' # Acme Orders project guidance - Run tests with `python -m pytest -q`. - Keep currency behavior explicit; never assume USD. - Add or update focused tests with every behavior change. - Read the smallest relevant set of files first, then expand when evidence requires it. # Compact instructions Preserve decisions, changed files, failing test output, and the next action. EOFThis overwrites the existing
CLAUDE.md. -
In the first Terminal, at the Claude prompt
>, start a fresh Claude Code session: Run/clear.- This command will be explained a bit more in the next Step!
-
Then run
/context.Confirm the
CLAUDE.mdfile appears under memory or project instructions. -
Then prompt Claude with
What test command and currency rule apply in this project?.It will return the
pythoncommand from theCLAUDE.mdyou just created.
-
-
Challenge
Manage accumulated context
/compactsummarizes older conversation history to free up context while preserving important information and continuity. Use it when continuing related work and you need to retain key decisions, changes, or findings. Avoid it when your next task is unrelated and the previous context is no longer useful./clearstarts a fresh session context and resets the session’s usage totals. Use it when switching to an unrelated task or when stale context is causing confusion. Avoid it when you still need detailed decisions, evidence, or progress from the current task.You can guide compaction, for example:
/compact Preserve changed files, test results, currency decisions, and unresolved questions.Compact proactively while the session still contains a clean, coherent record. If you wait until the session is crowded, stale, or contradictory, the resulting summary has less signal to preserve. Proactive compaction has a cost—it rebuilds context and may reduce detail—but that cost is justified when it protects continuity for substantial related work.
-
Challenge
Model and effort are separate cost levers
This lab doesn't support swapping models, but it is important to note that model and effort swapping for taskes does help optimize token usage.
Context management controls how much information is carried. Model selection controls which model processes it. Effort controls how much reasoning work a supported model applies. These levers are related but not interchangeable.
- Use a lighter or lower-cost model for routine, bounded edits and simple explanations.
- Use a stronger model for architecture, ambiguous multi-file diagnosis, or decisions requiring judgment.
- Lower effort for mechanical, well-specified work.
- Raise effort when the task benefits from deeper reasoning, trade-off analysis, or verification.
You can also launch with an explicit effort level, such as
claude --effort low, if supported by your model and organization settings. -
Challenge
Subagent deligations
A subagent works in its own context and returns a focused result to the main session. This is useful when exploration would flood the main context with search results, logs, or many file contents that will not be needed afterward. It is not free: the subagent consumes its own tokens and adds coordination overhead. Subagents can be delegated using a prompt such as
Use a subagent to inspect the repository for every place currency is created, assumed, converted, formatted, or tested. Do not edit files. Return only: (1) the relevant file paths, (2) the currency data flow, and (3) any defect or missing test.Use a subagent when
- the exploration is large but the main session needs only a summary;
- the task can be clearly bounded; and
- isolation is worth the extra request and coordination.
Keep work in the main session when the evidence must remain available for repeated reasoning or when the task is too small to justify delegation.
-
Challenge
Agent scope restriction and quality tradeoff
Token efficiency is not the same as task quality. A narrow prompt can omit a contract, caller, test, configuration, or architectural rule that changes the correct solution. For multi-file work, start focused and expand based on evidence rather than enforcing an arbitrary one-file limit.
In this project, asking Claude to fix
src/receipts.pywithout seeingmodels.py,pricing.py,service.py,docs/architecture.md, ortests/could produce a plausible local patch while missing the actual currency contract (specified in theCLAUDE.mdfile). Use the smallest resource level that preserves the evidence and reasoning the task requires.-
New unrelated task: Use
/clearto remove stale context from future turns. -
Continuing a long related task: Use
/compactwith preservation instructions to retain important decisions while freeing context space. -
Simple bounded edit: Target only the relevant files and use a lighter model or lower effort level to reduce token usage.
-
Ambiguous cross-file failure: Expand the relevant context and use a stronger model or higher effort when the solution requires cross-file understanding and careful judgment.
-
Noisy one-time exploration: Use a focused subagent to return a concise summary without filling the main session with unnecessary details.
-
Poor answer after excessive restriction: Restore the necessary files or increase the model or effort level. The additional cost is justified when more evidence or reasoning improves the result.
A proactive compaction is justified when continued work needs continuity but the session is becoming crowded. A model upgrade is justified when the expected cost of a weak decision, missed dependency, or repeated retry exceeds the extra model cost.
-
About the author
Real skill practice before real-world application
Hands-on Labs are real environments created by industry experts to help you learn. These environments help you gain knowledge and experience, practice without compromising your system, test without risk, destroy without fear, and let you learn from your mistakes. Hands-on Labs: practice your skills before delivering in the real world.
Learn by doing
Engage hands-on with the tools and technologies you’re learning. You pick the skill, we provide the credentials and environment.
Follow your guide
All labs have detailed instructions and objectives, guiding you through the learning process and ensuring you understand every step.
Turn time into mastery
On average, you retain 75% more of your learning if you take time to practice. Hands-on labs set you up for success to make those skills stick.