Skip to main content
Create your own

Building a Streaming Chat API with OpenAI Compatibility

Introduction

Welcome to Module 12. You've successfully completed the capstone project, building and benchmarking a custom LLM inference engine. This was a deep dive into the core mechanics of token generation. Now, we shift our focus from the engine to the service. A powerful engine is useless if applications can't communicate with it effectively. This module, Production Systems: Concurrency and Stateful Services, is about building the robust, user-facing layers that transform your engine into a production-grade product.

In this first lesson, we'll tackle the most crucial layer: the API. Your learning outcome is to implement an OpenAI-compatible chat completions API with streaming support on top of your custom inference engine.

Why this specific API? Because it is the de-facto industry standard. By conforming to this standard, your engine instantly becomes compatible with a vast ecosystem of tools, libraries, and applications. We will use FastAPI to build this API wrapper, covering both standard and streaming responses.

Let's visualize where this new component fits. The API server acts as the front door to your engine.

LLM Engine Architecture Overview
This diagram from the vLLM documentation shows how an OpenAI-compatible API Server (the component we are building today) wraps the core `AsyncLLMEngine` to handle incoming web requests.

Why an OpenAI-Compatible API?

Choosing an API standard is a significant architectural decision. Adopting the OpenAI specification provides immediate, powerful advantages that accelerate development and integration.

To understand the industry consensus around this, let's watch a short clip.

Running a High Throughput OpenAI-Compatible vLLM Inference Server on Modal

In this video, a developer from Modal explains why they chose to implement an OpenAI-compatible server for their vLLM deployment. His reasoning highlights the practical benefits you'll gain by doing the same.

Watch the segment from 01:39 to 03:30. Pay attention to the discussion on how the interface for LLMs has evolved beyond simple text-in/text-out and how the industry has coalesced around the OpenAI API standard.

As the video explains, the key benefits are:

  • Seamless Integration: Your engine will work out-of-the-box with any tool designed to work with OpenAI models, including the official openai Python client, LangChain, LlamaIndex, and countless others.
  • Developer Familiarity: Developers are already accustomed to the OpenAI API's structure, making your service easier to adopt.
  • Ecosystem Support: You leverage a rich ecosystem for testing, validation, and even monetization that is built around this standard.

The Anatomy of the Chat Completions API

To build a compatible API, you must precisely implement its data contracts. We will define these using Pydantic models, which integrate natively with FastAPI.

Let's examine the specification, covering the request, the non-streaming response, and the streaming response format.

OpenAI-Compatible API Server: Production-Ready LLM Serving

This article, 'OpenAI-Compatible API Server: Production-Ready LLM Serving,' provides a superb breakdown of the required Pydantic models. We will use this as our reference specification.

Read the sections 'Architecture Overview' and 'API Protocol Implementation'. Focus on the Pydantic models provided: ChatCompletionRequest: The input structure. ChatCompletionResponse: The output for non-streaming requests. ChatCompletionStreamResponse: The structure of each chunk in a streaming response. Familiarize yourself with the key fields like messages, choices, delta, and usage. We will be implementing these directly.

Let's summarize the key components:

1. The Request: ChatCompletionRequest

This model defines what a client sends to your server. The most important fields are:

  • model: A string identifying the model to use (e.g., "my-custom-phi3").
  • messages: A list of message objects, each with a role ("system", "user", or "assistant") and content.
  • stream: A boolean that determines whether to receive the response as a single JSON object or a stream of events.
  • Other parameters like max_tokens, temperature, top_p control the generation process.
2. The Non-Streaming Response: ChatCompletionResponse

When stream=False, the server waits for the entire generation to complete and then sends a single JSON response.

OpenAI-Compatible Chat Completions API Response Example
This image shows a complete, non-streaming JSON response from an OpenAI-compatible API. Note the `choices` array containing the full assistant message and the `usage` object with token counts.

This response includes the full generated message in choices[0].message.content and a finish_reason (e.g., "stop" or "length").

3. The Streaming Response: ChatCompletionStreamResponse

When stream=True, the server uses the Server-Sent Events (SSE) protocol. Instead of one big response, it sends a series of small messages as tokens are generated. Each message is a "chunk" and has a specific format:

  • Each chunk is a JSON object.
  • The first chunk usually specifies the role (e.g., delta: {"role": "assistant"}).
  • Subsequent chunks contain the generated tokens in the delta.content field.
  • The final chunk includes the finish_reason.
  • Each message is prefixed with data: and ends with \n\n.
  • The stream is terminated by a special message: data: [DONE]\n\n.

This protocol is what enables the "typewriter" effect in applications like ChatGPT, drastically improving the perceived latency.

Implementation with FastAPI and Your Custom Engine

Now, let's write the code. You will create a FastAPI application that wraps your custom engine and exposes the /v1/chat/completions endpoint.

We'll assume your capstone engine has an interface that can be adapted to provide an asynchronous generator for tokens, similar to engine.generate_stream_async(prompt, params).

Step 1: Project Setup
  1. Create a new Python file for your API server (e.g., api_server.py).
  2. Install fastapi, uvicorn, and openai: pip install fastapi uvicorn openai.
  3. Import necessary libraries and your engine components.
  4. Copy the Pydantic models for ChatCompletionRequest, ChatCompletionResponse, etc., from the LINK article into your file. They are your data contracts.
Step 2: Create the FastAPI Endpoint

Define the basic FastAPI application and the endpoint. The endpoint function will take a ChatCompletionRequest object as input.

from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
import uvicorn




# Assume your pydantic models from the article are defined here
# Assume your engine is imported and instantiated
# from my_engine import CustomEngine
# engine = CustomEngine(...)

app = FastAPI()

@app.post("/v1/chat/completions")
async def create_chat_completion(request: ChatCompletionRequest):



    # We will implement the logic here
    # For now, a placeholder
    if request.stream:
        return StreamingResponse(...)
    else:
        return ...
Step 3: Implement the Streaming Logic

Implementing streaming correctly is the most involved part. You'll need an async def generator function that yields each chunk in the correct SSE format.

The video below gives a clear, practical demonstration of using FastAPI's StreamingResponse with a generator.

Streaming LLM Responses with FastAPI

This tutorial, 'Streaming LLM Responses with FastAPI', provides an excellent, concise walkthrough of the core mechanics. It shows how to connect a generator function to a StreamingResponse object.

Watch from 01:57 to 08:47. First, the video explains the non-streaming approach. Then, it dives into the streaming implementation. Focus on three key things: The use of llm.stream() to get a generator of chunks. The async def generator function that yields the content. How this generator is passed to StreamingResponse.

Now, let's apply this pattern to generate OpenAI-compatible chunks. Your create_chat_completion function will check the request.stream flag and delegate to a generator function.

import json
import time
import uuid




# ... (inside your api_server.py)

async def stream_generator(request: ChatCompletionRequest):
    """
    An async generator that calls the engine and formats the output
    into OpenAI-compatible Server-Sent Events.
    """



    # Unique ID for the request
    request_id = f"chatcmpl-{uuid.uuid4()}"
    



    # 1. Send the first chunk with the role
    initial_chunk = {
        "id": request_id,
        "object": "chat.completion.chunk",
        "created": int(time.time()),
        "model": request.model,
        "choices": [{
            "index": 0,
            "delta": {"role": "assistant"},
            "finish_reason": None
        }]
    }
    yield f"data: {json.dumps(initial_chunk)}\n\n"




    # 2. Call your engine's streaming method and yield token chunks
    # This assumes your engine has a method that yields tokens.
    # You will need to adapt this to your specific engine's interface.
    # We'll also assume your engine's `apply_chat_template` method
    # converts the `messages` list to a single prompt string.
    prompt = your_tokenizer.apply_chat_template(request.messages, tokenize=False)
    
    final_finish_reason = "stop" # Default
    try:
        async for output in engine.generate_stream_async(prompt, request):



            # The 'output' here would be the raw token string from your engine.
            # You might need to handle token usage counts here as well.
            
            token_chunk = {
                "id": request_id,
                "object": "chat.completion.chunk",
                "created": int(time.time()),
                "model": request.model,
                "choices": [{
                    "index": 0,
                    "delta": {"content": output.text},
                    "finish_reason": None # Will be set in the final chunk
                }]
            }
            yield f"data: {json.dumps(token_chunk)}\n\n"
            final_finish_reason = output.finish_reason # Update based on engine output

    except Exception as e:
        print(f"Error during generation: {e}")

```grasp
{
  "type": "exercise",
  "id": "3d6a3e64-6c47-4d43-877f-7b3970ac6608"
}
    # Handle errors if necessary
    final_finish_reason = "error"





# 3. Send the final chunk with the finish reason
final_chunk = {
    "id": request_id,
    "object": "chat.completion.chunk",
    "created": int(time.time()),
    "model": request.model,
    "choices": [{
        "index": 0,
        "delta": {},
        "finish_reason": final_finish_reason
    }]
}
yield f"data: {json.dumps(final_chunk)}\n\n"




# 4. Send the [DONE] signal
yield "data: [DONE]\n\n"

Now, update your main endpoint function

@app.post("/v1/chat/completions")
async def create_chat_completion(request: ChatCompletionRequest):

# Here you would convert the request into parameters for your engine
# (e.g., sampling_params from request.temperature, request.top_p, etc.)

if request.stream:
    return StreamingResponse(stream_generator(request), media_type="text/event-stream")
else:



    # Implementation for non-streaming
    # 1. Call your engine's non-streaming generate method
    # prompt = your_tokenizer.apply_chat_template(...)
    # result = await engine.generate_async(prompt, request)
    



    # 2. Format the result into the ChatCompletionResponse pydantic model
    # response = ChatCompletionResponse(...)
    



    # return response
    raise HTTPException(status_code=501, detail="Non-streaming not implemented yet.")
**Note:** The code above is a template. You'll need to integrate it with your actual engine's method for generating tokens and handling chat templates. The key is the structure of the `stream_generator` function and its use of `yield` to produce correctly formatted SSE messages.


```grasp
{
  "type": "exercise",
  "id": "a333eb17-6c9f-4ff6-927b-995840db71a0"
}
Step 4: Testing Your API

Once your server is running (uvicorn api_server:app --reload), you can test it using curl or the openai Python client.

Using curl for a streaming request:

curl -N http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "your-model-name",
    "messages": [
      {"role": "user", "content": "Write a short poem about building an API."}
    ],
    "max_tokens": 100,
    "stream": true
  }'

The -N flag disables buffering in curl, so you'll see the chunks arrive one by one.

Using the openai Python client:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="dummy-key" # The key is not used but is required by the client
)

stream = client.chat.completions.create(
    model="your-model-name",
    messages=[{"role": "user", "content": "Tell me a joke about computers."}],
    stream=True,
)

for chunk in stream:
    if chunk.choices[0].delta.content is not None:
        print(chunk.choices[0].delta.content, end="")

print()

This client-side code demonstrates the power of compatibility—it works with your custom server exactly as it would with OpenAI's.

Conclusion

In this lesson, you've wrapped your powerful custom engine in a production-standard API. This is a critical step in turning a research project into a usable service.

Key Takeaways:

  • Industry Standards Matter: Adopting the OpenAI API specification makes your engine instantly compatible with a massive ecosystem of tools and familiar to developers.
  • Pydantic Defines the Contract: Using Pydantic models with FastAPI is the canonical way to enforce API data structures for requests and responses.
  • Streaming is Essential for UX: For chat applications, streaming responses via Server-Sent Events (SSE) is non-negotiable for good user experience. FastAPI's StreamingResponse and async generators provide a powerful pattern for implementing this.
  • The SSE Protocol is Specific: A compliant streaming response requires correctly formatted chunks (data: <json>\n\n) and a final [DONE] message.

Preview of the Next Lesson:
Our current API is stateless. Each request is treated as a new, independent conversation. However, real chat applications are stateful, requiring the model to remember previous turns. In the next lesson, you will design and implement an in-memory session store for managing concurrent, multi-turn conversation states. This will involve handling the growing list of messages in each request and exploring strategies for managing conversation context efficiently.

Can't find a good explanation? Sign up and we'll make it for you

Sign up