Featured resource
2026 Tech Forecast
2026 Tech Forecast

1,500+ tech insiders, business leaders, and Pluralsight Authors share their predictions on what’s shifting fastest and how to stay ahead.

Download the forecast
  • Lab
    • Libraries: If you want this lab, consider one of these libraries.
    • AI
Labs

Token Management and Cost Optimization with Claude Code

In this lab, you will learn to manage Claude Code’s token usage, session context, and model selection to balance cost, performance, and output quality during agentic coding workflows.

Lab platform
Lab Info
Level
Intermediate
Last updated
Oct 07, 2026
Duration
30m

Contact sales

By clicking submit, you agree to our Privacy Policy and Terms of Use, and consent to receive marketing emails from Pluralsight.
Table of Contents
  1. Challenge

    Introduction

    In this lab, you’ll explore how token usage and session context affect the cost, speed, and quality of Claude Code workflows. Through hands-on activities, you’ll practice monitoring usage, managing context, selecting appropriate models and effort levels, and balancing token efficiency with task quality.

    Claude has already been set up, and some code and configuration files were provided, to help guide you through the foundations of token management and cost optimization.

  2. Challenge

    Token and session overview

    Tokens and session context

    A token is a unit of text processed by a language model. A token may be a word, part of a word, punctuation, or whitespace. Token counts therefore do not map one-to-one to word counts.

    For each turn, Claude Code builds an active context that can include the following:

    • system and project instructions;
    • the conversation history that has not been cleared or summarized;
    • file contents and tool results brought into the session;
    • your current prompt; and
    • space for Claude's response and reasoning.

    Conceptually, this active context accompanies each new request. Prompt caching can make repeated prefixes cheaper and faster, and Claude Code may clear old tool results or compact history, but retained context still occupies the working window. This is why a short new question late in a long session can consume far more input tokens than the same question in a fresh session.

    One way to understand this at a high-level is as follows:

    turn cost ≈ retained input context + new input + generated output + reasoning

    The exact bill depends on the model, authentication plan, caching, provider, and organization pricing. 1. At the Claude Code prompt, enter /context

    1. Then examine the usage with /usage

    2. Press esc to close the view.

    3. Then prompt Claude with the following: Read README.md and explain the project in two sentences. Do not inspect other files.

    4. Run /context and /usage again, and note the change in session metrics.

    /usage shows session token details and a locally computed cost estimate for API usage. Subscription users may see plan-usage information instead; an estimated dollar value is not necessarily their billed amount. ### Growing context matters

    As active context grows:

    • Cost can rise: more input must be processed on subsequent turns, although cached tokens may be priced differently.
    • Responses can slow: larger prompts take more time to process.
    • Quality can decline: stale or conflicting details can distract the model, and the most relevant evidence can become harder to prioritize.
    • Compaction approaches: Claude Code warns as the auto-compact threshold nears and eventually summarizes older history to free space.

    Warning signs include slower turns, /context showing little free space, an auto-compact warning, repeated confusion with old requirements, or rising input totals for otherwise similar tasks. A context warning is not the same as a subscription or spending limit.

  3. Challenge

    Minimizing processing tokens

    Specific prompts reduce irrelevant exploration and unnecessary file content.

    1. Run /usage, and grab a baseline for the comparisons you will do next.

    2. Compare these approaches: Read the entire repository and tell me how tax works. and Read src/pricing.py and tests/test_pricing.py. Explain how tax is calculated.

    3. Using /usage to see how rapidly the token limit raises from a broad search versus a narrowed seach

    Note: broad requests can still be benefitical in cases of architecture discovery, but should be used sparingly.

  4. Challenge

    Preserve recurring context with CLAUDE.md

    A project-root CLAUDE.md contains recurring project instructions that Claude Code loads at the start of sessions. It avoids repeatedly typing conventions and can survive compaction because the root file is re-read. It still consumes context, so keep it concise and reserve it for durable information rather than temporary task notes.

    1. To the right of the Terminal, click the +, select New tab, then click Terminal.

    2. In the new Terminal, at the workspace $ prompt, run cd claude_token_lab.

      Then run the following:

      cat > CLAUDE.md <<'EOF'
      # Acme Orders project guidance
      
      - Run tests with `python -m pytest -q`.
      - Keep currency behavior explicit; never assume USD.
      - Add or update focused tests with every behavior change.
      - Read the smallest relevant set of files first, then expand when evidence requires it.
      
      # Compact instructions
      Preserve decisions, changed files, failing test output, and the next action.
      EOF
      

      This overwrites the existing CLAUDE.md.

    3. In the first Terminal, at the Claude prompt >, start a fresh Claude Code session: Run /clear.

      • This command will be explained a bit more in the next Step!
    4. Then run /context.

      Confirm the CLAUDE.md file appears under memory or project instructions.

    5. Then prompt Claude with What test command and currency rule apply in this project?.

      It will return the python command from the CLAUDE.md you just created.

  5. Challenge

    Manage accumulated context

    /compact summarizes older conversation history to free up context while preserving important information and continuity. Use it when continuing related work and you need to retain key decisions, changes, or findings. Avoid it when your next task is unrelated and the previous context is no longer useful.

    /clear starts a fresh session context and resets the session’s usage totals. Use it when switching to an unrelated task or when stale context is causing confusion. Avoid it when you still need detailed decisions, evidence, or progress from the current task.

    You can guide compaction, for example:

    /compact Preserve changed files, test results, currency decisions, and unresolved questions.

    Compact proactively while the session still contains a clean, coherent record. If you wait until the session is crowded, stale, or contradictory, the resulting summary has less signal to preserve. Proactive compaction has a cost—it rebuilds context and may reduce detail—but that cost is justified when it protects continuity for substantial related work.

  6. Challenge

    Model and effort are separate cost levers

    This lab doesn't support swapping models, but it is important to note that model and effort swapping for taskes does help optimize token usage.

    Context management controls how much information is carried. Model selection controls which model processes it. Effort controls how much reasoning work a supported model applies. These levers are related but not interchangeable.

    • Use a lighter or lower-cost model for routine, bounded edits and simple explanations.
    • Use a stronger model for architecture, ambiguous multi-file diagnosis, or decisions requiring judgment.
    • Lower effort for mechanical, well-specified work.
    • Raise effort when the task benefits from deeper reasoning, trade-off analysis, or verification.

    You can also launch with an explicit effort level, such as claude --effort low, if supported by your model and organization settings.

  7. Challenge

    Subagent deligations

    A subagent works in its own context and returns a focused result to the main session. This is useful when exploration would flood the main context with search results, logs, or many file contents that will not be needed afterward. It is not free: the subagent consumes its own tokens and adds coordination overhead. Subagents can be delegated using a prompt such as

    Use a subagent to inspect the repository for every place currency is created, assumed, converted, formatted, or tested. Do not edit files. Return only: (1) the relevant file paths, (2) the currency data flow, and (3) any defect or missing test.

    Use a subagent when

    • the exploration is large but the main session needs only a summary;
    • the task can be clearly bounded; and
    • isolation is worth the extra request and coordination.

    Keep work in the main session when the evidence must remain available for repeated reasoning or when the task is too small to justify delegation.

  8. Challenge

    Agent scope restriction and quality tradeoff

    Token efficiency is not the same as task quality. A narrow prompt can omit a contract, caller, test, configuration, or architectural rule that changes the correct solution. For multi-file work, start focused and expand based on evidence rather than enforcing an arbitrary one-file limit.

    In this project, asking Claude to fix src/receipts.py without seeing models.py, pricing.py, service.py, docs/architecture.md, or tests/ could produce a plausible local patch while missing the actual currency contract (specified in the CLAUDE.md file). Use the smallest resource level that preserves the evidence and reasoning the task requires.

    • New unrelated task: Use /clear to remove stale context from future turns.

    • Continuing a long related task: Use /compact with preservation instructions to retain important decisions while freeing context space.

    • Simple bounded edit: Target only the relevant files and use a lighter model or lower effort level to reduce token usage.

    • Ambiguous cross-file failure: Expand the relevant context and use a stronger model or higher effort when the solution requires cross-file understanding and careful judgment.

    • Noisy one-time exploration: Use a focused subagent to return a concise summary without filling the main session with unnecessary details.

    • Poor answer after excessive restriction: Restore the necessary files or increase the model or effort level. The additional cost is justified when more evidence or reasoning improves the result.

    A proactive compaction is justified when continued work needs continuity but the session is becoming crowded. A model upgrade is justified when the expected cost of a weak decision, missed dependency, or repeated retry exceeds the extra model cost.

About the author

I am, Josh Meier, an avid explorer of ideas an a lifelong learner. I have a background in AI with a focus in generative AI. I am passionate about AI and the ethics surrounding its use and creation and have honed my skills in generative AI models, ethics and applications and thrive to improve in my understanding of these models.

Real skill practice before real-world application

Hands-on Labs are real environments created by industry experts to help you learn. These environments help you gain knowledge and experience, practice without compromising your system, test without risk, destroy without fear, and let you learn from your mistakes. Hands-on Labs: practice your skills before delivering in the real world.

Learn by doing

Engage hands-on with the tools and technologies you’re learning. You pick the skill, we provide the credentials and environment.

Follow your guide

All labs have detailed instructions and objectives, guiding you through the learning process and ensuring you understand every step.

Turn time into mastery

On average, you retain 75% more of your learning if you take time to practice. Hands-on labs set you up for success to make those skills stick.

Get started with Pluralsight