Thariq explains how prompt caching powers long‑running agents in Claude Code, describing the ordering of static prompts and tools to maximise cache hits. He shares practical techniques such as using system‑reminder messages, deferring tool loading, and safe compaction to avoid costly cache misses.
Highlights
Prompt caching reuses prior computation via prefix matching, reducing latency and cost for agents.
Static system prompts and tools should be placed before dynamic session context to maximize shared prefixes.
Changing static prompt content or tool definitions causes cache misses and higher expenses.
Techniques include embedding updates in messages, using stub tools, and exact parent‑prefix matching.
auto-generated
Thariq · via X (formerly Twitter)
Context
Audience
AI engineers and product developers building agentic LLM systems
LLM prompt engineeringcache optimization techniquesClaude Code documentationAnthropic API
Discover Similar Content
ngrok.com
Prompt caching: 10x cheaper LLM tokens, but how? | ngrok blog
A far more detailed explanation of prompt caching than anyone asked for.
github.com
GitHub - Dicklesworthstone/post_compact_reminder: Claude Code hook that detects context compaction and injects a reminder to re-read AGENTS.md, preventing post-compaction rule amnesia in long sessions
Claude Code hook that detects context compaction and injects a reminder to re-read AGENTS.md, preventing post-compaction rule amnesia in long sessions...
engineering.atspotify.com
Portal by Spotify cut my Claude Code token usage by 90% | Spotify Engineering
Most of what an AI coding agent does for me isn’t thinking. It’s I/O.