mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
279 lines
11 KiB
Text
279 lines
11 KiB
Text
|
|
{
|
||
|
|
"cells": [
|
||
|
|
{
|
||
|
|
"cell_type": "markdown",
|
||
|
|
"metadata": {},
|
||
|
|
"source": [
|
||
|
|
"# Context Compression\n",
|
||
|
|
"\n",
|
||
|
|
"## What is it\n",
|
||
|
|
"\n",
|
||
|
|
"*Context Compression is the act of statistically reducing tool output size while preserving the information the LLM needs to answer the user's question.*\n",
|
||
|
|
"\n",
|
||
|
|
"## Why it helps\n",
|
||
|
|
"\n",
|
||
|
|
"* Avoids [Context Distraction](https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html): Verbose tool outputs dilute the signal. Compression removes filler words and redundant phrasing while keeping key facts, errors, and anomalies.\n",
|
||
|
|
"* **No extra LLM call required**: Unlike pruning (notebook 04) and summarization (notebook 05) which call GPT-4o-mini per tool result, compression runs locally using statistical and ML-based token analysis. Zero additional cost, lower latency.\n",
|
||
|
|
"\n",
|
||
|
|
"## Context Compression in Practice\n",
|
||
|
|
"\n",
|
||
|
|
"[Headroom](https://github.com/chopratejas/headroom) is an open-source context optimization library that provides multi-algorithm compression. It auto-detects content type (JSON, code, logs, text) and routes to the optimal compressor:\n",
|
||
|
|
"\n",
|
||
|
|
"- **SmartCrusher**: Statistically analyzes JSON arrays \u2014 keeps errors, anomalies, and query-relevant items\n",
|
||
|
|
"- **Kompress**: ModernBERT token classifier \u2014 removes redundant tokens from text while preserving meaning\n",
|
||
|
|
"- **CodeCompressor**: AST-aware compression for source code\n",
|
||
|
|
"\n",
|
||
|
|
"When items are highly diverse (like RAG retriever chunks), Headroom keeps all items and compresses the text *within* each one \u2014 no information is dropped.\n",
|
||
|
|
"\n",
|
||
|
|
"## Context Compression in LangGraph\n",
|
||
|
|
"\n",
|
||
|
|
"We'll replace the LLM-based pruning/summarization step with a local compression call. The agent structure is identical to notebooks 04 and 05 \u2014 only the tool processing node changes."
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"# Install headroom (one-time)\n",
|
||
|
|
"# !pip install \"headroom-ai[all]\""
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from langchain_community.document_loaders import WebBaseLoader\n",
|
||
|
|
"\n",
|
||
|
|
"urls = [\n",
|
||
|
|
" \"https://lilianweng.github.io/posts/2025-05-01-thinking/\",\n",
|
||
|
|
" \"https://lilianweng.github.io/posts/2024-11-28-reward-hacking/\",\n",
|
||
|
|
" \"https://lilianweng.github.io/posts/2024-07-07-hallucination/\",\n",
|
||
|
|
" \"https://lilianweng.github.io/posts/2024-04-12-diffusion-video/\",\n",
|
||
|
|
"]\n",
|
||
|
|
"\n",
|
||
|
|
"docs = [WebBaseLoader(url).load() for url in urls]"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from langchain_text_splitters import RecursiveCharacterTextSplitter\n",
|
||
|
|
"\n",
|
||
|
|
"docs_list = [item for sublist in docs for item in sublist]\n",
|
||
|
|
"\n",
|
||
|
|
"text_splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(\n",
|
||
|
|
" chunk_size=3000, chunk_overlap=50\n",
|
||
|
|
")\n",
|
||
|
|
"doc_splits = text_splitter.split_documents(docs_list)"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from langchain.embeddings import init_embeddings\n",
|
||
|
|
"from langchain_core.vectorstores import InMemoryVectorStore\n",
|
||
|
|
"\n",
|
||
|
|
"embeddings = init_embeddings(\"openai:text-embedding-3-small\")\n",
|
||
|
|
"vectorstore = InMemoryVectorStore.from_documents(documents=doc_splits, embedding=embeddings)\n",
|
||
|
|
"retriever = vectorstore.as_retriever()"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from langchain.tools.retriever import create_retriever_tool\n",
|
||
|
|
"from rich.console import Console\n",
|
||
|
|
"from rich.pretty import pprint\n",
|
||
|
|
"\n",
|
||
|
|
"console = Console()\n",
|
||
|
|
"\n",
|
||
|
|
"retriever_tool = create_retriever_tool(\n",
|
||
|
|
" retriever,\n",
|
||
|
|
" \"retrieve_blog_posts\",\n",
|
||
|
|
" \"Search and return information about Lilian Weng blog posts.\",\n",
|
||
|
|
")\n",
|
||
|
|
"\n",
|
||
|
|
"result = retriever_tool.invoke({\"query\": \"types of reward hacking\"})\n",
|
||
|
|
"console.print(\"[bold green]Retriever Tool Results:[/bold green]\")\n",
|
||
|
|
"pprint(result)"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from langchain.chat_models import init_chat_model\n",
|
||
|
|
"\n",
|
||
|
|
"llm = init_chat_model(\"anthropic:claude-sonnet-4-20250514\", temperature=0)\n",
|
||
|
|
"\n",
|
||
|
|
"tools = [retriever_tool]\n",
|
||
|
|
"tools_by_name = {tool.name: tool for tool in tools}\n",
|
||
|
|
"\n",
|
||
|
|
"llm_with_tools = llm.bind_tools(tools)"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from typing import Literal\n",
|
||
|
|
"\n",
|
||
|
|
"from IPython.display import Image, display\n",
|
||
|
|
"from langchain_core.messages import SystemMessage, ToolMessage\n",
|
||
|
|
"from langgraph.graph import END, START, MessagesState, StateGraph\n",
|
||
|
|
"\n",
|
||
|
|
"from headroom import compress\n",
|
||
|
|
"\n",
|
||
|
|
"\n",
|
||
|
|
"class State(MessagesState):\n",
|
||
|
|
" \"\"\"Extended state that includes a summary field for context compression.\"\"\"\n",
|
||
|
|
"\n",
|
||
|
|
" summary: str\n",
|
||
|
|
"\n",
|
||
|
|
"\n",
|
||
|
|
"rag_prompt = \"\"\"You are a helpful assistant tasked with retrieving information from a series of technical blog posts by Lilian Weng.\n",
|
||
|
|
"Clarify the scope of research with the user before using your retrieval tool to gather context. Reflect on any context you fetch, and\n",
|
||
|
|
"proceed until you have sufficient context to answer the user's research request.\"\"\"\n",
|
||
|
|
"\n",
|
||
|
|
"\n",
|
||
|
|
"def llm_call(state: State) -> dict:\n",
|
||
|
|
" \"\"\"Execute LLM call with system prompt and message history.\"\"\"\n",
|
||
|
|
" messages = [SystemMessage(content=rag_prompt)] + state[\"messages\"]\n",
|
||
|
|
" response = llm_with_tools.invoke(messages)\n",
|
||
|
|
" return {\"messages\": [response]}\n",
|
||
|
|
"\n",
|
||
|
|
"\n",
|
||
|
|
"def should_continue(state: State) -> Literal[\"tool_node_with_compression\", \"__end__\"]:\n",
|
||
|
|
" \"\"\"Decide if we should continue the loop or stop.\"\"\"\n",
|
||
|
|
" messages = state[\"messages\"]\n",
|
||
|
|
" last_message = messages[-1]\n",
|
||
|
|
" if last_message.tool_calls:\n",
|
||
|
|
" return \"tool_node_with_compression\"\n",
|
||
|
|
" return END\n",
|
||
|
|
"\n",
|
||
|
|
"\n",
|
||
|
|
"def tool_node_with_compression(state: State):\n",
|
||
|
|
" \"\"\"Execute tool calls and compress results with Headroom.\n",
|
||
|
|
"\n",
|
||
|
|
" Instead of calling GPT-4o-mini to prune or summarize (notebooks 04, 05),\n",
|
||
|
|
" we use Headroom's compress() \u2014 no LLM call, no extra cost.\n",
|
||
|
|
"\n",
|
||
|
|
" Headroom auto-detects content type and applies the right compressor:\n",
|
||
|
|
" - JSON arrays \u2192 SmartCrusher (statistical, keeps anomalies + query-relevant items)\n",
|
||
|
|
" - Plain text \u2192 Kompress (ModernBERT token compression)\n",
|
||
|
|
" - Code \u2192 CodeCompressor (AST-aware)\n",
|
||
|
|
"\n",
|
||
|
|
" For diverse retriever results (each chunk is unique), Headroom keeps ALL\n",
|
||
|
|
" items and compresses the text within each one.\n",
|
||
|
|
" \"\"\"\n",
|
||
|
|
" result = []\n",
|
||
|
|
" for tool_call in state[\"messages\"][-1].tool_calls:\n",
|
||
|
|
" tool = tools_by_name[tool_call[\"name\"]]\n",
|
||
|
|
" observation = tool.invoke(tool_call[\"args\"])\n",
|
||
|
|
"\n",
|
||
|
|
" # Build a minimal message list so Headroom can extract the user query\n",
|
||
|
|
" # for relevance-aware compression (keeps chunks matching the question).\n",
|
||
|
|
" user_query = state[\"messages\"][0].content if state[\"messages\"] else \"\"\n",
|
||
|
|
" temp_messages = [\n",
|
||
|
|
" {\"role\": \"user\", \"content\": user_query},\n",
|
||
|
|
" {\"role\": \"tool\", \"content\": observation, \"tool_call_id\": tool_call[\"id\"]},\n",
|
||
|
|
" ]\n",
|
||
|
|
"\n",
|
||
|
|
" compressed = compress(temp_messages, model=\"claude-sonnet-4-20250514\")\n",
|
||
|
|
" compressed_content = compressed.messages[-1][\"content\"]\n",
|
||
|
|
"\n",
|
||
|
|
" result.append(ToolMessage(content=compressed_content, tool_call_id=tool_call[\"id\"]))\n",
|
||
|
|
"\n",
|
||
|
|
" return {\"messages\": result}\n",
|
||
|
|
"\n",
|
||
|
|
"\n",
|
||
|
|
"# Build workflow\n",
|
||
|
|
"agent_builder = StateGraph(State)\n",
|
||
|
|
"\n",
|
||
|
|
"agent_builder.add_node(\"llm_call\", llm_call)\n",
|
||
|
|
"agent_builder.add_node(\"tool_node_with_compression\", tool_node_with_compression)\n",
|
||
|
|
"\n",
|
||
|
|
"agent_builder.add_edge(START, \"llm_call\")\n",
|
||
|
|
"agent_builder.add_conditional_edges(\n",
|
||
|
|
" \"llm_call\",\n",
|
||
|
|
" should_continue,\n",
|
||
|
|
" {\n",
|
||
|
|
" \"tool_node_with_compression\": \"tool_node_with_compression\",\n",
|
||
|
|
" END: END,\n",
|
||
|
|
" },\n",
|
||
|
|
")\n",
|
||
|
|
"agent_builder.add_edge(\"tool_node_with_compression\", \"llm_call\")\n",
|
||
|
|
"\n",
|
||
|
|
"agent = agent_builder.compile()\n",
|
||
|
|
"\n",
|
||
|
|
"display(Image(agent.get_graph(xray=True).draw_mermaid_png()))"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "code",
|
||
|
|
"execution_count": null,
|
||
|
|
"metadata": {},
|
||
|
|
"outputs": [],
|
||
|
|
"source": [
|
||
|
|
"from utils import format_messages\n",
|
||
|
|
"\n",
|
||
|
|
"query = \"What are the types of reward hacking discussed in the blogs?\"\n",
|
||
|
|
"result = agent.invoke({\"messages\": [{\"role\": \"user\", \"content\": query}]})\n",
|
||
|
|
"format_messages(result[\"messages\"])"
|
||
|
|
]
|
||
|
|
},
|
||
|
|
{
|
||
|
|
"cell_type": "markdown",
|
||
|
|
"metadata": {},
|
||
|
|
"source": [
|
||
|
|
"## How it compares\n",
|
||
|
|
"\n",
|
||
|
|
"| Technique | Notebook | Token Reduction | Extra LLM Call | Extra Cost |\n",
|
||
|
|
"|-----------|----------|----------------|----------------|------------|\n",
|
||
|
|
"| RAG Baseline | 01 | \u2014 | No | $0 |\n",
|
||
|
|
"| Context Pruning | 04 | ~56% | Yes (GPT-4o-mini) | ~$0.003/call |\n",
|
||
|
|
"| Context Summarization | 05 | ~68% | Yes (GPT-4o-mini) | ~$0.003/call |\n",
|
||
|
|
"| **Context Compression** | **07** | **~30-40%** | **No** | **$0** |\n",
|
||
|
|
"\n",
|
||
|
|
"Key differences:\n",
|
||
|
|
"\n",
|
||
|
|
"- **No LLM call**: Pruning and summarization call GPT-4o-mini per tool result. Compression runs locally.\n",
|
||
|
|
"- **No information loss**: For diverse retriever results (each chunk is unique), Headroom keeps ALL items and compresses text within each one. Pruning removes entire chunks; summarization rewrites them.\n",
|
||
|
|
"- **Reversible**: Headroom's CCR (Compress-Cache-Retrieve) stores originals. The LLM can call `headroom_retrieve` to get full uncompressed content if it needs more detail.\n",
|
||
|
|
"- **Content-aware**: Different content types get different treatment. JSON arrays \u2192 statistical analysis. Plain text \u2192 ML token compression. Code \u2192 AST-aware compression.\n",
|
||
|
|
"\n",
|
||
|
|
"The trade-off: pruning and summarization can achieve higher compression (56-68%) because they use an LLM to judge relevance. Compression achieves 30-40% without any LLM call \u2014 making it faster and free."
|
||
|
|
]
|
||
|
|
}
|
||
|
|
],
|
||
|
|
"metadata": {
|
||
|
|
"kernelspec": {
|
||
|
|
"display_name": "Python 3 (ipykernel)",
|
||
|
|
"language": "python",
|
||
|
|
"name": "python3"
|
||
|
|
},
|
||
|
|
"language_info": {
|
||
|
|
"name": "python",
|
||
|
|
"version": "3.11.0"
|
||
|
|
}
|
||
|
|
},
|
||
|
|
"nbformat": 4,
|
||
|
|
"nbformat_minor": 4
|
||
|
|
}
|