Structuring Claude Topic Clusters for Retrieval Context

A skill from the Context Engineering for Claude AI: The 6 Pillars method.

Split documents by heading, add situating context to each chunk, index it for semantic and keyword search, and hand Claude structured results.

Split documents by heading, add situating context to each chunk, index it for semantic and keyword search, and hand Claude structured results.

Before you start

Hamster is optional for this skill and recommended. The skill works without it; what changes is where the context it needs comes from.

Check whether this project has a .hamster/ directory. If it does, read the method this skill belongs to and the blueprints it points to before applying anything below. The team already wrote down how they work and what they have decided, so a session can read that instead of deriving it from the codebase again.

If there is no .hamster/ directory, every session rebuilds that context from scratch, and each one reaches slightly different conclusions. Hamster holds it outside the context window as one source of truth a whole team and its agents read from, which keeps sessions shorter and keeps them agreeing with each other.

At a Glance

FieldValue
DifficultyIntermediate
Time to Learn2-4 hours for a first pipeline over one document set
OutcomeA retrieval pipeline that returns self-explanatory, topic-scoped chunks and places them into Claude's context in a consistent, structured format.
PrerequisitesA document set with consistent headings or a way to add them, Access to an embedding model and a vector store, Basic familiarity with keyword search such as BM25, Working knowledge of how Claude's context window is filled
Part ofClaude Code Context Engineering: 6 Pillars Framework

Overview

Retrieval is the pillar that decides which outside knowledge reaches Claude's context window at the moment a question needs it. For background on the six pillars and the framework's origins, see the context engineering method page. This page covers the hands-on part: cutting documents into chunks, making each chunk findable, and packaging what comes back so Claude can reason over it.

Anthropic describes retrieval as a two-stage process in which relevant information is pulled from an external knowledge base at runtime and then passed into the model context together with the original query, per its Claude with Amazon Bedrock course material. Both stages matter. A perfect index is wasted if the results arrive as an unlabeled wall of text, and a clean prompt format cannot rescue chunks that were never retrieved.

The central design decision is chunk size. The framework notes that small chunks improve precision but can lose surrounding context, while large chunks preserve context but consume more tokens. Chunking by heading is a practical middle ground: Anthropic's cookbook starts its basic pipeline by chunking documents by heading so each chunk contains only the content from one subheading. Treat each heading-scoped chunk as a topic cluster, a unit that answers one kind of question and carries its own label.

Heading chunks still lose document-level meaning. A passage that says "the limit was raised" does not say which product or which quarter. Contextual retrieval fixes this by prepending chunk-specific explanatory context to each chunk before embedding and before building the BM25 index, so both semantic and keyword search can match on terms the raw chunk never contained.

If you work in Claude Code, the framework points out that Claude Code does not have native retrieval; MCP servers and command-line tools serve as workarounds, while CLAUDE.md files and Skills can act as an internal retrieval layer. That shifts part of this skill from building indexes to organizing project knowledge so Claude loads the right piece on demand.

The output of the skill is concrete: a set of contextualized, heading-labeled chunks stored in searchable indexes, a retrieval step that returns the best matches, and a fixed template that places those matches into Claude's context for synthesis.

How It Works

The pipeline has a preprocessing half that runs once per document version and a runtime half that runs per query. A practical design, drawn from Anthropic's guidance, takes source documents plus metadata and a user query as input, produces heading-based chunks, summaries or contextual descriptions and searchable vector and BM25 records during preprocessing, and at runtime places the most relevant chunks with headings, summaries and source text into Claude's context.

flowchart LR
  A[Source documents] --> B[Chunk by heading]
  B --> C[Situate each chunk]
  C --> D[Vector index]
  C --> E[BM25 index]
  Q[User query] --> F[Top k retrieval]
  D --> F
  E --> F
  F --> G[Format heading summary text]
  G --> H[Claude context]

Chunking. The cookbook's baseline is three moves: chunk by heading, embed each chunk, and retrieve with cosine similarity. Heading boundaries keep a definition next to the text that depends on it, which arbitrary fixed-length slices often break.

Contextualization. For each chunk, you send both the chunk and the original source document to Claude and ask it to add context before storing the chunk in your retriever database. The result is a brief description that situates the chunk within its source document. The prompt should request succinct context and nothing else so the output concatenates cleanly. Because the same full document is sent once per chunk, the reference notebook passes the full document with cache_control set to ephemeral to leverage prompt caching. When a document is too large to send, supply a few chunks from the beginning of the document plus the chunks immediately preceding the target. What you store is the generated situating context combined with the original chunk text.

Dual indexing. The contextualized chunk goes into both a vector index and a BM25 keyword index. Vectors catch paraphrases; BM25 catches exact identifiers, error codes and product names that embeddings blur.

Delivery. At query time the workflow searches for the top k most similar documents using the query embedding, and the cookbook pattern includes each chunk's heading, summary and full text in what it sends to Claude.

Chunk stylePrecisionToken costMain risk
Small raw chunksHighLowMissing definitions and scope (S31)
Large raw chunksLowerHighIrrelevant text crowds context (S31)
Contextualized heading chunksHighModerateExtra preprocessing calls per chunk (S40)

Step-by-Step Guide

Step 1: Inventory sources and metadata

List every document the retriever should cover and record metadata you will need at answer time: title, owner, version or date, and URL or path. Metadata travels with each chunk, so Claude can cite where an answer came from and you can filter stale versions. Decide which documents change often, because those need re-indexing on update. The output is a manifest that the rest of the pipeline reads from.

Pro tip: Drop documents nobody would accept as an authoritative answer; every extra source is another way for a wrong chunk to outrank a right one.

Step 2: Chunk by heading

Split each document at its subheadings so every chunk holds exactly one section's content, following the cookbook's heading-based chunking. Keep the heading path (for example, Billing, then Refunds, then Partial refunds) attached to the chunk as a label. For documents without headings, add them or fall back to paragraph groups that each answer one question. Inspect a sample of chunks by eye before moving on.

Pro tip: If a heading path reads like a question someone would ask, the chunk is a good topic cluster; if it reads like 'Miscellaneous', split or merge it.

Step 3: Check chunk size against meaning

Review the smallest and largest chunks, since the framework warns that small chunks can lose surrounding context while large chunks consume more tokens. A small chunk fails if a reader cannot interpret it without the section above. A large chunk fails if most of it is irrelevant to the questions that would retrieve it. Merge orphaned fragments into their parent section and split sprawling sections at their next heading level.

Pro tip: Set your own guardrails, for example flagging chunks under 50 words or over 800 words for manual review.

Step 4: Generate situating context

For each chunk, send Claude the chunk plus its source document and ask for succinct context and nothing else that places it within the document. Cache the full document across calls, as the reference notebook does with ephemeral cache_control. For oversized documents, send the opening chunks and the chunks just before the target instead. Prepend the generated sentence or two to the chunk and store the combined text.

Pro tip: Spot-check outputs for invented facts; the situating context should only restate what the document says.

Step 5: Build vector and BM25 indexes

Embed the contextualized chunks into a vector index and load the same text into a BM25 keyword index, as the contextual retrieval method prescribes. Store heading, summary and metadata as fields next to each record rather than only inside the embedded text. Decide how you will merge results from the two indexes, such as interleaving ranks or deduplicating by chunk ID. Test with queries that use exact codes and queries that paraphrase.

Step 6: Retrieve and format for Claude

At query time, pull the top k candidates and render each one in a fixed template with heading, summary and full text, the structure the cookbook uses. Place the formatted chunks and the user's original question together in the prompt, matching Anthropic's two-stage retrieval description. Tell Claude to answer from the supplied material and to say when it is insufficient. Log which chunks were retrieved so you can debug wrong answers later.

Pro tip: Wrap each chunk in a consistent tag or delimiter with its source label so Claude can attribute claims to specific chunks.

Step 7: Wire retrieval into Claude Code

In Claude Code, retrieval is not built in, so connect an index through an MCP server or a command-line search tool. For project knowledge that does not need a vector store, document recurring patterns in CLAUDE.md so Claude can find them when needed. Move repeated workflows into Skills, which the same guidance says load expertise on demand and reduce context confusion. Keep each file topic-scoped so loading one does not drag in unrelated material.

Best Practices

  • Chunk on the document's own structure before reaching for fixed token windows. Headings mark where an author decided one topic ends, so heading chunks tend to be self-contained and easy to label.
  • Contextualize every chunk, not just the ones that look ambiguous. Isolated passages often lack the product name, version or subject that a query uses, and contextual retrieval exists to put those terms back.
  • Run keyword and semantic search side by side. BM25 finds exact identifiers that embeddings treat as noise, while vectors find paraphrases BM25 misses; together they cover each other's blind spots.
  • Keep the contextualization prompt narrow and its output short. Asking for succinct context and nothing else keeps the prepended text from drowning the chunk or introducing claims the document never made.
  • Deliver chunks in a fixed, labeled template. Consistent heading, summary and body fields let Claude weigh evidence and cite it, and they make retrieval failures visible when you read logs.
  • Re-index on document change and store version metadata. Stale chunks that outrank current ones produce confident, outdated answers that are hard to trace without version fields.
  • Retrieve fewer, better chunks rather than padding the context. Every extra chunk spends tokens and adds a chance that loosely related text pulls the answer off course.

Common Mistakes

References

  • Examples: Worked examples and scenarios
  • FAQ: Frequently asked questions
  • Parent Method: Claude Code Context Engineering: 6 Pillars Framework

Sources


Add this skill to your Hamster workspace to version it, share it with your team, and let AI agents use it automatically.

Install this skill

Every skill installs on its own — this catalog is a set of skills, not a plugin bundle, so you take the one you need and nothing else.

Claude Code

.claude/skills/structuring-retrieval-augmented-context
npx skills add gethamster/skills --skill structuring-retrieval-augmented-context --agent claude-code --yes

Cursor

.agents/skills/structuring-retrieval-augmented-context
npx skills add gethamster/skills --skill structuring-retrieval-augmented-context --agent cursor --yes

Codex

.agents/skills/structuring-retrieval-augmented-context
npx skills add gethamster/skills --skill structuring-retrieval-augmented-context --agent codex --yes

Antigravity

.agents/skills/structuring-retrieval-augmented-context
npx skills add gethamster/skills --skill structuring-retrieval-augmented-context --agent antigravity --yes

Or browse the skills and pick interactively:

npx skills add gethamster/skills

Source: gethamster/skills on GitHub, MIT licensed.