Case study 2 · Reference architecture
Hybrid Streaming RAG
Retrieval-augmented generation on streaming data that takes personal data, auditability and partial failure seriously.
The problem
Retrieval-augmented generation on streaming data must handle personal data, auditability and partial failures. Most demos ignore all three.
The approach
- 01Personal data never leaves the ingress raw. Recognised personal data is replaced with versioned HMAC tokens before chunking, indexing, inference, events or audit writes.
- 02Traceable requests. Every request gets server-issued request and correlation IDs and a W3C trace context. These travel in events and Kafka headers, and caller-supplied values are discarded.
- 03Defined failure behaviour. If a downstream system fails, the workflow ends in a defined state, completed or compensated, with a matching compensating event.
- 04Immutable audit through object-lock storage on S3, Azure or GCS.
- 05Operational controls: bulkheads, rate limits, and Prometheus metrics with bounded cardinality.
- 06Formal architecture documents: a TOGAF Architecture Definition and C4 models from Level 1 to Level 4.
flowchart LR C["Client"] --> A["API ingress<br/>request_id · correlation_id<br/>W3C traceparent"] A --> T["PII replaced with<br/>versioned HMAC tokens"] T --> K["Chunk and index"] --> V["Vector store adapter<br/>Chroma or OpenSearch"] T --> M["Inference / risk scorer<br/>adapter: Ollama"] A --> E["Kafka events<br/>with correlation headers"] A --> U["Immutable audit<br/>object lock: S3 · Azure · GCS"] A --> L["Lifecycle<br/>COMPLETED or COMPENSATED"]
Honest scope
The README itself says this is a bounded compensating-transaction design, not full event sourcing.Repository →