Fast Answer online serving architecture
The User and Chat Service communicate bidirectionally. The Chat Service calls a stateless Fast Answer Service over gRPC with a 200 millisecond timeout. Inside the service, the query is normalized, an in-memory Catalog is queried, and an in-process Decision model is evaluated. Optional user features may be retrieved under a strict timeout. On fallback, the Chat Service calls the LLM Generation service.
Fast Answer Service
Horizontally scaled stateless replica
Stateless
Kubernetes
CLIENT
User
ORCHESTRATOR
Chat Service
First-turn requests only
FALLBACK PATH
LLM Generation
Normal response path
REQUEST PATH
Query Normalization
Exact-match key
PROCESS-LOCAL
In-Memory Catalog
QA pair
Catalog features
Immutable version
IN-PROCESS
Decision Model
Lightweight model
Decision threshold
OPTIONAL REMOTE CALL
User Feature Store
gRPC
200 ms timeout
Fallback