Introduction
Welcome to Module 12. You've successfully completed the capstone project, building and benchmarking a custom LLM inference engine. This was a deep dive into the core mechanics of token generation. Now, we shift our focus from the engine to the service. A powerful engine is useless if applications can't communicate with it effectively. This module, Production Systems: Concurrency and Stateful Services, is about building the robust, user-facing layers that transform your engine into a production-grade product.
In this first lesson, we'll tackle the most crucial layer: the API. Your learning outcome is to implement an OpenAI-compatible chat completions API with streaming support on top of your custom inference engine.
Why this specific API? Because it is the de-facto industry standard. By conforming to this standard, your engine instantly becomes compatible with a vast ecosystem of tools, libraries, and applications. We will use FastAPI to build this API wrapper, covering both standard and streaming responses.
Let's visualize where this new component fits. The API server acts as the front door to your engine.

Why an OpenAI-Compatible API?
Choosing an API standard is a significant architectural decision. Adopting the OpenAI specification provides immediate, powerful advantages that accelerate development and integration.
To understand the industry consensus around this, let's watch a short clip.
Running a High Throughput OpenAI-Compatible vLLM Inference Server on Modal
In this video, a developer from Modal explains why they chose to implement an OpenAI-compatible server for their vLLM deployment. His reasoning highlights the practical benefits you'll gain by doing the same.
Watch the segment from 01:39 to 03:30. Pay attention to the discussion on how the interface for LLMs has evolved beyond simple text-in/text-out and how the industry has coalesced around the OpenAI API standard.
As the video explains, the key benefits are:
- Seamless Integration: Your engine will work out-of-the-box with any tool designed to work with OpenAI models, including the official
openaiPython client, LangChain, LlamaIndex, and countless others. - Developer Familiarity: Developers are already accustomed to the OpenAI API's structure, making your service easier to adopt.
- Ecosystem Support: You leverage a rich ecosystem for testing, validation, and even monetization that is built around this standard.
The Anatomy of the Chat Completions API
To build a compatible API, you must precisely implement its data contracts. We will define these using Pydantic models, which integrate natively with FastAPI.
Let's examine the specification, covering the request, the non-streaming response, and the streaming response format.
OpenAI-Compatible API Server: Production-Ready LLM Serving
This article, 'OpenAI-Compatible API Server: Production-Ready LLM Serving,' provides a superb breakdown of the required Pydantic models. We will use this as our reference specification.
Read the sections 'Architecture Overview' and 'API Protocol Implementation'. Focus on the Pydantic models provided: ChatCompletionRequest: The input structure. ChatCompletionResponse: The output for non-streaming requests. ChatCompletionStreamResponse: The structure of each chunk in a streaming response. Familiarize yourself with the key fields like messages, choices, delta, and usage. We will be implementing these directly.
Let's summarize the key components:
1. The Request: ChatCompletionRequest
This model defines what a client sends to your server. The most important fields are:
model: A string identifying the model to use (e.g.,"my-custom-phi3").messages: A list of message objects, each with arole("system","user", or"assistant") andcontent.stream: A boolean that determines whether to receive the response as a single JSON object or a stream of events.- Other parameters like
max_tokens,temperature,top_pcontrol the generation process.
2. The Non-Streaming Response: ChatCompletionResponse
When stream=False, the server waits for the entire generation to complete and then sends a single JSON response.

This response includes the full generated message in choices[0].message.content and a finish_reason (e.g., "stop" or "length").
3. The Streaming Response: ChatCompletionStreamResponse
When stream=True, the server uses the Server-Sent Events (SSE) protocol. Instead of one big response, it sends a series of small messages as tokens are generated. Each message is a "chunk" and has a specific format:
- Each chunk is a JSON object.
- The first chunk usually specifies the role (e.g.,
delta: {"role": "assistant"}). - Subsequent chunks contain the generated tokens in the
delta.contentfield. - The final chunk includes the
finish_reason. - Each message is prefixed with
data:and ends with\n\n. - The stream is terminated by a special message:
data: [DONE]\n\n.
This protocol is what enables the "typewriter" effect in applications like ChatGPT, drastically improving the perceived latency.
Implementation with FastAPI and Your Custom Engine
Now, let's write the code. You will create a FastAPI application that wraps your custom engine and exposes the /v1/chat/completions endpoint.
We'll assume your capstone engine has an interface that can be adapted to provide an asynchronous generator for tokens, similar to engine.generate_stream_async(prompt, params).
Step 1: Project Setup
- Create a new Python file for your API server (e.g.,
api_server.py). - Install
fastapi,uvicorn, andopenai:pip install fastapi uvicorn openai. - Import necessary libraries and your engine components.
- Copy the Pydantic models for
ChatCompletionRequest,ChatCompletionResponse, etc., from the LINK article into your file. They are your data contracts.
Step 2: Create the FastAPI Endpoint
Define the basic FastAPI application and the endpoint. The endpoint function will take a ChatCompletionRequest object as input.
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
import uvicorn
# Assume your pydantic models from the article are defined here
# Assume your engine is imported and instantiated
# from my_engine import CustomEngine
# engine = CustomEngine(...)
app = FastAPI()
@app.post("/v1/chat/completions")
async def create_chat_completion(request: ChatCompletionRequest):
# We will implement the logic here
# For now, a placeholder
if request.stream:
return StreamingResponse(...)
else:
return ...
Step 3: Implement the Streaming Logic
Implementing streaming correctly is the most involved part. You'll need an async def generator function that yields each chunk in the correct SSE format.
The video below gives a clear, practical demonstration of using FastAPI's StreamingResponse with a generator.
Streaming LLM Responses with FastAPI
This tutorial, 'Streaming LLM Responses with FastAPI', provides an excellent, concise walkthrough of the core mechanics. It shows how to connect a generator function to a StreamingResponse object.
Watch from 01:57 to 08:47. First, the video explains the non-streaming approach. Then, it dives into the streaming implementation. Focus on three key things: The use of llm.stream() to get a generator of chunks. The async def generator function that yields the content. How this generator is passed to StreamingResponse.
Now, let's apply this pattern to generate OpenAI-compatible chunks. Your create_chat_completion function will check the request.stream flag and delegate to a generator function.
import json
import time
import uuid
# ... (inside your api_server.py)
async def stream_generator(request: ChatCompletionRequest):
"""
An async generator that calls the engine and formats the output
into OpenAI-compatible Server-Sent Events.
"""
# Unique ID for the request
request_id = f"chatcmpl-{uuid.uuid4()}"
# 1. Send the first chunk with the role
initial_chunk = {
"id": request_id,
"object": "chat.completion.chunk",
"created": int(time.time()),
"model": request.model,
"choices": [{
"index": 0,
"delta": {"role": "assistant"},
"finish_reason": None
}]
}
yield f"data: {json.dumps(initial_chunk)}\n\n"
# 2. Call your engine's streaming method and yield token chunks
# This assumes your engine has a method that yields tokens.
# You will need to adapt this to your specific engine's interface.
# We'll also assume your engine's `apply_chat_template` method
# converts the `messages` list to a single prompt string.
prompt = your_tokenizer.apply_chat_template(request.messages, tokenize=False)
final_finish_reason = "stop" # Default
try:
async for output in engine.generate_stream_async(prompt, request):
# The 'output' here would be the raw token string from your engine.
# You might need to handle token usage counts here as well.
token_chunk = {
"id": request_id,
"object": "chat.completion.chunk",
"created": int(time.time()),
"model": request.model,
"choices": [{
"index": 0,
"delta": {"content": output.text},
"finish_reason": None # Will be set in the final chunk
}]
}
yield f"data: {json.dumps(token_chunk)}\n\n"
final_finish_reason = output.finish_reason # Update based on engine output
except Exception as e:
print(f"Error during generation: {e}")
```grasp
{
"type": "exercise",
"id": "3d6a3e64-6c47-4d43-877f-7b3970ac6608"
}
# Handle errors if necessary
final_finish_reason = "error"
# 3. Send the final chunk with the finish reason
final_chunk = {
"id": request_id,
"object": "chat.completion.chunk",
"created": int(time.time()),
"model": request.model,
"choices": [{
"index": 0,
"delta": {},
"finish_reason": final_finish_reason
}]
}
yield f"data: {json.dumps(final_chunk)}\n\n"
# 4. Send the [DONE] signal
yield "data: [DONE]\n\n"
Now, update your main endpoint function
@app.post("/v1/chat/completions")
async def create_chat_completion(request: ChatCompletionRequest):
# Here you would convert the request into parameters for your engine
# (e.g., sampling_params from request.temperature, request.top_p, etc.)
if request.stream:
return StreamingResponse(stream_generator(request), media_type="text/event-stream")
else:
# Implementation for non-streaming
# 1. Call your engine's non-streaming generate method
# prompt = your_tokenizer.apply_chat_template(...)
# result = await engine.generate_async(prompt, request)
# 2. Format the result into the ChatCompletionResponse pydantic model
# response = ChatCompletionResponse(...)
# return response
raise HTTPException(status_code=501, detail="Non-streaming not implemented yet.")
**Note:** The code above is a template. You'll need to integrate it with your actual engine's method for generating tokens and handling chat templates. The key is the structure of the `stream_generator` function and its use of `yield` to produce correctly formatted SSE messages.
```grasp
{
"type": "exercise",
"id": "a333eb17-6c9f-4ff6-927b-995840db71a0"
}
Step 4: Testing Your API
Once your server is running (uvicorn api_server:app --reload), you can test it using curl or the openai Python client.
Using curl for a streaming request:
curl -N http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "your-model-name",
"messages": [
{"role": "user", "content": "Write a short poem about building an API."}
],
"max_tokens": 100,
"stream": true
}'
The -N flag disables buffering in curl, so you'll see the chunks arrive one by one.
Using the openai Python client:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="dummy-key" # The key is not used but is required by the client
)
stream = client.chat.completions.create(
model="your-model-name",
messages=[{"role": "user", "content": "Tell me a joke about computers."}],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="")
print()
This client-side code demonstrates the power of compatibility—it works with your custom server exactly as it would with OpenAI's.
Conclusion
In this lesson, you've wrapped your powerful custom engine in a production-standard API. This is a critical step in turning a research project into a usable service.
Key Takeaways:
- Industry Standards Matter: Adopting the OpenAI API specification makes your engine instantly compatible with a massive ecosystem of tools and familiar to developers.
- Pydantic Defines the Contract: Using Pydantic models with FastAPI is the canonical way to enforce API data structures for requests and responses.
- Streaming is Essential for UX: For chat applications, streaming responses via Server-Sent Events (SSE) is non-negotiable for good user experience. FastAPI's
StreamingResponseand async generators provide a powerful pattern for implementing this. - The SSE Protocol is Specific: A compliant streaming response requires correctly formatted chunks (
data: <json>\n\n) and a final[DONE]message.
Preview of the Next Lesson:
Our current API is stateless. Each request is treated as a new, independent conversation. However, real chat applications are stateful, requiring the model to remember previous turns. In the next lesson, you will design and implement an in-memory session store for managing concurrent, multi-turn conversation states. This will involve handling the growing list of messages in each request and exploring strategies for managing conversation context efficiently.