Skip to main content
Create your own

Optimizing Stateful Multi-Turn Tool-Calling with KV Cache Reuse

Introduction

In the previous lesson, you successfully built a stateful chat service by implementing an in-memory session store. You made the crucial connection between the application-level concept of conversation history and the system-level optimization of KV cache reuse, particularly through mechanisms like SGLang's RadixAttention.

Today, we'll expand on this stateful capability to handle a more complex and powerful interaction pattern: tool-calling. An agent that can only talk is limited; an agent that can act by calling tools (APIs, functions, etc.) is far more useful.

Your learning outcome for this lesson is to implement stateful multi-turn tool-calling, focusing on KV cache reuse across tool-call rounds and its impact on the scheduler.

This introduces a fascinating systems challenge. When an LLM decides to call a tool, it pauses its generation to wait for the tool's result. During this pause, what should the inference engine do with the valuable GPU memory occupied by the request's KV cache? This is not just a memory management question; it's a fundamental scheduling problem that pits the latency of a single request against the throughput of the entire system. We will explore the trade-offs and design a principled scheduling policy to manage them.

The Application-Level Tool-Calling Loop

Let's first understand the logical flow of a single turn involving a tool call. Unlike a simple chat response, this is a multi-step process:

  1. User to LLM: The user submits a prompt (e.g., "What's the weather in Paris?").
  2. LLM to Tool: The LLM, having been trained for function calling, doesn't answer directly. Instead, it generates a structured output indicating a tool needs to be called (e.g., {'tool_name': 'get_weather', 'arguments': {'location': 'Paris'}}).
  3. Server-Side Execution: Your API server parses this output, recognizes it as a tool call, and executes the corresponding function (e.g., calls an external weather API).
  4. Tool to LLM: The server takes the tool's output (e.g., {'temperature': '15°C', 'condition': 'Cloudy'}), formats it into a message, and appends it to the conversation history.
  5. LLM to User: The server sends the entire updated history back to the LLM, which now has the context of the tool's result. The LLM then generates the final, user-facing response (e.g., "The weather in Paris is 15°C and cloudy.").

This entire sequence constitutes a single "turn" from the user's perspective, managed within the same session_id you implemented in the last lesson.

The System-Level Challenge: Managing the "Pause"

From a systems engineering perspective, the critical moment is between steps 2 and 4—the "pause" when your server is executing the tool. During this time, the LLM request is idle, but its KV cache, which represents the entire conversation history up to the tool call, is holding onto precious GPU VRAM.

The inference engine's scheduler faces a dilemma:

  • Policy 1: Evict the KV Cache. The scheduler can free the KV cache associated with the waiting request, allowing other requests to use the GPU. When the tool result is ready, the original request is re-queued.
    • Pro: Maximizes GPU utilization and system throughput.
    • Con: Incurs a significant latency penalty. The entire conversation history must be re-processed (a full prefill) to reconstruct the KV cache before generation can resume.
  • Policy 2: Pin the KV Cache. The scheduler can "pin" the KV cache in GPU memory, preserving it until the tool call completes.
    • Pro: Extremely low latency for the returning request, as no re-computation is needed.
    • Con: Terrible for system throughput. The pinned memory is unusable by other requests, effectively reducing the GPU's capacity. If the tool is slow, the GPU is underutilized.

Neither of these naive policies is optimal. Eviction is too slow for interactive applications, while pinning is too wasteful for a high-throughput service. We need a more intelligent approach.

Prefix Caching: The Core Mechanism for Reuse

Before we design a better policy, let's revisit the mechanism that makes reuse possible. The concept is often called prefix caching. In a tool-calling scenario, the entire conversation history before the tool's output is a prefix. When the tool's output arrives, we want to append it and only compute the KV tensors for this new information, reusing the cache for the prefix.

To solidify this idea, let's watch a short video that explains the general concept of prompt caching.

What is Prompt Caching? Optimize LLM Latency with AI Transformers

This IBM Technology video provides a clear, high-level explanation of prompt caching. It directly addresses caching common prefixes like system prompts, few-shot examples, and critically for us, tool definitions and conversation history.

Watch from 01:47 to 07:52. Pay close attention to how it distinguishes between static, cacheable content (the prefix) and dynamic content (the new user question). This is directly analogous to our tool-calling loop, where the chat history is the prefix and the tool's output is the new dynamic content.

This video confirms that structuring our prompts to have long, static prefixes is key to enabling efficient caching. In multi-turn chat and tool-use, the conversation history is that prefix.

Designing a Principled Scheduling Policy

The core of today's lesson is to move beyond the naive "always evict" or "always pin" policies. The decision to retain a KV cache should be an economic one, balancing the cost of occupying memory against the benefit of avoiding re-computation.

This is precisely the problem addressed by recent research in LLM serving systems. We will now study the approach taken by Continuum, a serving system designed specifically for these kinds of agentic workloads.

Continuum's key idea is to introduce a Time-to-Live (TTL) for a pinned KV cache. Instead of pinning indefinitely, the scheduler pins the cache for a calculated duration. If the tool call finishes within the TTL, the request gets the full low-latency benefit. If it takes too long, the cache is evicted, freeing the GPU for other work, thus preventing system stalls.

This transforms the problem from a binary choice (evict/pin) into an optimization problem: What is the optimal TTL?

To understand how Continuum answers this, you'll read sections from the original paper. As a researcher, you'll appreciate seeing how these system-level problems are formalized and solved.

Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

The paper 'Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live' directly tackles the scheduling challenge of agentic workloads. We will focus on the motivation, the core algorithm, and the utility model used to make scheduling decisions.

Please read the following sections to understand Continuum's approach: Section 1 (Introduction) & Section 3.2 (Failure Analysis for Prior Work): These sections set up the problem. Focus on why 'end-of-turn eviction' fails for agents and the concept of accumulating 'per-turn queueing delay'. This is the 'why' behind the need for a better scheduler. Section 4 (Continuum Scheduling Algorithm): Read the introduction to this section and the part discussing the TTL trade-off (Figure 6). This introduces the core concept of using TTL. Section 4.1 (Utility Model): This is the most important part. Study the formulas for 'Cost Estimation' and 'Benefit Estimation'. You don't need to memorize them, but understand the components: the cost is the opportunity cost of blocked memory, while the benefit is the sum of avoided recomputation (Cost_reload) and avoided queueing delay (Cost_ordering). Note the innovative 'memoryfulness factor' used to model queueing delay. Section 4.2 (Setting the TTL Value): Skim this section to see how the utility model is used to find the optimal TTL that maximizes the expected net benefit, using the historical distribution of tool-call durations. This reading will provide a robust mental model of how a modern inference scheduler handles stateful, multi-turn agent execution.

The Continuum paper provides the final piece of the puzzle. We have the application logic (the tool-use loop), the memory mechanism (PagedAttention/RadixAttention for prefix sharing), and now the scheduling policy (TTL-based retention) that makes the whole system efficient.

Visualizing Stateful Management with RadixAttention

Let's tie this back to the RadixAttention mechanism we discussed in the last lesson. A tool-calling conversation creates a chain of dependencies in the KV cache, which a radix tree is perfectly suited to manage. When a tool call happens, the scheduler might evict the leaf node (the last turn) to save memory, but the parent nodes (the rest of the history) remain cached. Continuum's TTL policy is what would guide this eviction decision.

KV Cache Reuse and Eviction in Multi-Turn Conversations
This figure, which we saw before, illustrates how the radix tree evolves. A tool call represents a pause at one of the nodes (e.g., node 'a' in step 2). Continuum's scheduler would decide how long to keep node 'a' and its KV cache in memory. If the TTL expires, it might get evicted (like node 'c' in step 5) to make space for another request.

The sharing pattern for a multi-turn conversation, whether it includes tool calls or not, is the same. The scheduler's intelligence lies in managing the lifecycle of these cache entries.

KV Cache Sharing Examples in LLM Interaction Patterns
As shown in example (c) for Multi-Turn Chat, the 'Chat History' is the shareable prefix. In a tool-calling context, this history would include not just user/assistant messages but also tool call requests and tool results.

High-Level Implementation Sketch

Fully implementing Continuum is beyond a 60-minute lesson, but we can sketch out how you would modify your custom engine's scheduler to incorporate this logic.

Your scheduler, which currently manages a queue of requests, would need to be enhanced:

class EnhancedScheduler:
    def __init__(self):
        self.waiting_queue = []
        self.running_batch = []



        # Maps session_id to a pinned KV cache and its expiry timestamp
        self.pinned_caches = {} 

    async def schedule_step(self):



        # 1. Evict expired caches
        current_time = asyncio.get_event_loop().time()
        for session_id, pin_info in list(self.pinned_caches.items()):
            if current_time > pin_info['expires_at']:
                print(f"TTL expired for session {session_id}. Evicting KV cache.")
                self.engine.free_kv_cache(pin_info['cache_handle'])
                del self.pinned_caches[session_id]




        # 2. Build the next batch
        # Prioritize requests that have a corresponding pinned cache
        # (This is the 'TTL-aware priority' from Continuum)
        next_batch = []



        # ... logic to select requests from waiting_queue, prioritizing those in pinned_caches




        # 3. Run the batch
        # results = await self.engine.execute(next_batch)




        # 4. Process results and decide on pinning
        for result in results:



            # Assume result has session_id, output_text, and finish_reason
            if result.finish_reason == 'tool_call':



                # Parse the tool call
                tool_info = parse_tool_call(result.output_text)




                # Calculate TTL based on Continuum's utility model (simplified here)
                ttl = self.calculate_ttl(tool_info, result.kv_cache_size)
                
                if ttl > 0:
                    print(f"Pinning KV cache for session {result.session_id} with TTL {ttl}s.")
                    self.pinned_caches[result.session_id] = {
                        'cache_handle': result.kv_cache_handle,
                        'expires_at': asyncio.get_event_loop().time() + ttl
                    }
                else:



                    # If TTL is 0, evict immediately
                    self.engine.free_kv_cache(result.kv_cache_handle)
            else:



                # Normal finish, free the cache
                self.engine.free_kv_cache(result.kv_cache_handle)

    def calculate_ttl(self, tool_info, cache_size):



        # Placeholder for the complex utility model from the Continuum paper.
        # A simple heuristic could be: if tool is known to be fast, return a short TTL.
        # For a production system, you'd implement the full cost/benefit analysis.
        if tool_info['name'] in ['fast_local_lookup', 'simple_math']:
            return 2.0  # seconds
        else:
            return 0.5 # A very short TTL for unknown or slow tools

This pseudocode illustrates the key additions to a scheduler: a dictionary for pinned caches, logic to evict expired caches, and a decision point after each generation step to either pin or free the KV cache based on a calculated TTL.

Conclusion

In this lesson, you have tackled one of the most important and complex aspects of serving modern, agentic LLMs. You've gone beyond simple chat to understand the systems implications of tool use.

Key Takeaways:

  • Tool-calling creates an idle "pause" in the generation process, which poses a significant challenge for KV cache management.
  • Naive policies of always-evict or always-pin are suboptimal, forcing a harsh trade-off between latency and throughput.
  • The solution is a principled scheduling policy that makes an economic decision about retaining the KV cache.
  • Continuum's TTL-based retention provides a robust framework for this decision, balancing the benefit of cache reuse against the opportunity cost of occupying GPU memory.
  • This scheduling policy works in concert with underlying memory management mechanisms like PagedAttention and RadixAttention to enable efficient, stateful, multi-turn agent execution.

Preview of the Next Lesson:
We have now gone deep into the internals of the inference engine and its scheduler to optimize for complex stateful workloads. In the next lesson, we will zoom out to the API layer and address another critical aspect of production serving: implementing token-bucket rate limiting at the API gateway level. This will shift our focus from optimizing individual request performance to ensuring the stability and fairness of the service as a whole.

Can't find a good explanation? Sign up and we'll make it for you

Sign up