[ChatStream] Starting the Web Server (ASGI Server)
Hello from the Product Development Department at Qualiteg Inc.
In this article, we explain how to start a web server with ChatStream on board.
uvicorn (started from within the code)
Because ChatStream supports FastAPI/Starlette, it can be run on an ASGI server.
To define uvicorn within your code, implement it as follows:
def start_server():
uvicorn.run(app, host='localhost', port=9999)
def main():
start_server()
if __name__ == "__main__":
main()
Full source code
import torch
import uvicorn
from fastapi import FastAPI, Request
from fastersession import FasterSessionMiddleware, MemoryStore
from transformers import AutoTokenizer, AutoModelForCausalLM
from chatstream import ChatStream, ChatPromptTogetherRedPajamaINCITEChat as ChatPrompt
model_path = "togethercomputer/RedPajama-INCITE-Chat-3B-v1"
device = "cuda" # "cuda" / "cpu"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16)
model.to(device)
chat_stream = ChatStream(
num_of_concurrent_executions=2,
max_queue_size=5,
model=model,
tokenizer=tokenizer,
device=device,
chat_prompt_clazz=ChatPrompt,
)
app = FastAPI()
app.add_middleware(FasterSessionMiddleware,
secret_key="your-session-secret-key", # Key for cookie signature
store=MemoryStore(), # Store for session saving
http_only=True, # True: Cookie cannot be accessed from client-side scripts such as JavaScript
secure=True, # False: For local development env. True: For production. Requires Https
)
@app.post("/chat_stream")
async def stream_api(request: Request):
# handling FastAPI/Starlette's Request
response = await chat_stream.handle_chat_stream_request(request)
return response
@app.on_event("startup")
async def startup():
# start request queueing system
await chat_stream.start_queue_worker()
def start_server():
uvicorn.run(app, host='localhost', port=9999)
def main():
start_server()
if __name__ == "__main__":
main()
uvicorn (started externally)
Next, let's look at the pattern of starting uvicorn externally.
To start example_server_redpajama_simple.py in ./example as the server, run:
uvicorn example.web_server_redpajama_simple.py:app --host 0.0.0.0 --port 3000
This approach separates the server from the application, making it closer to a production setup.
uvicorn startup options
https://www.uvicorn.org/settings/
Source code
example_server_redpajama_simple.py
import torch
import uvicorn
from fastapi import FastAPI, Request
from fastersession import FasterSessionMiddleware, MemoryStore
from transformers import AutoTokenizer, AutoModelForCausalLM
from chatstream import ChatStream, ChatPromptTogetherRedPajamaINCITEChat as ChatPrompt
model_path = "togethercomputer/RedPajama-INCITE-Chat-3B-v1"
device = "cuda" # "cuda" / "cpu"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16)
model.to(device)
chat_stream = ChatStream(
num_of_concurrent_executions=2,
max_queue_size=5,
model=model,
tokenizer=tokenizer,
device=device,
chat_prompt_clazz=ChatPrompt,
)
app = FastAPI()
app.add_middleware(FasterSessionMiddleware,
secret_key="your-session-secret-key", # Key for cookie signature
store=MemoryStore(), # Store for session saving
http_only=True, # True: Cookie cannot be accessed from client-side scripts such as JavaScript
secure=True, # False: For local development env. True: For production. Requires Https
)
@app.post("/chat_stream")
async def stream_api(request: Request):
# handling FastAPI/Starlette's Request
response = await chat_stream.handle_chat_stream_request(request)
return response
@app.on_event("startup")
async def startup():
# start request queueing system
await chat_stream.start_queue_worker()
gunicorn
Next is the method using gunicorn. This is the most common approach for a production API server. By using the gunicorn instance started here as the API server and combining it with a web server such as Nginx acting as a reverse proxy, you can run it as a production system.
To start example_server_redpajama_simple.py in ./example as the server, run:
gunicorn example.web_server_redpajama_simple.py:app --workers 4 --worker-class uvicorn.workers.UvicornWorker --bind 0.0.0.0:3000
(Note: this does not work on Windows.)
Source code
example_server_redpajama_simple.py
import torch
from fastapi import FastAPI, Request
from fastersession import FasterSessionMiddleware, MemoryStore
from transformers import AutoTokenizer, AutoModelForCausalLM
from chatstream import ChatStream, ChatPromptTogetherRedPajamaINCITEChat as ChatPrompt
model_path = "togethercomputer/RedPajama-INCITE-Chat-3B-v1"
device = "cuda" # "cuda" / "cpu"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16)
model.to(device)
chat_stream = ChatStream(
num_of_concurrent_executions=2,
max_queue_size=5,
model=model,
tokenizer=tokenizer,
device=device,
chat_prompt_clazz=ChatPrompt,
)
app = FastAPI()
app.add_middleware(FasterSessionMiddleware,
secret_key="your-session-secret-key", # Key for cookie signature
store=MemoryStore(), # Store for session saving
http_only=True, # True: Cookie cannot be accessed from client-side scripts such as JavaScript
secure=True, # False: For local development env. True: For production. Requires Https
)
@app.post("/chat_stream")
async def stream_api(request: Request):
# handling FastAPI/Starlette's Request
response = await chat_stream.handle_chat_stream_request(request)
return response
@app.on_event("startup")
async def startup():
# start request queueing system
await chat_stream.start_queue_worker()
We hope this was helpful. Because ChatStream is implemented on top of FastAPI/Starlette, you can build a server for production using standard approaches.
That said, in a commercial environment it is rare to operate with just a single ChatStream API server. Each ChatStream API server is called a ChatStream node, and multiple ChatStream nodes form a cluster for each region. At Qualiteg, we recommend this kind of scale-out system configuration for handling load.