Why Events?
Our client Nexus Capital needed real-time portfolio updates across 500+ strategies processing 10M+ transactions daily. The legacy batch system delivered T+1 data—unacceptable for modern risk management.
Event-driven architecture was the answer. But the path from "let's use Kafka" to "10M events/day with exactly-once semantics" taught us hard lessons.
Lesson 1: Start with the Event Contract
Events are your API. Treat them like one.
{
"eventId": "uuid-v7",
"eventType": "portfolio.position.updated",
"version": "2.1.0",
"timestamp": "2024-11-15T10:30:00.000Z",
"source": "portfolio-service",
"data": {
"portfolioId": "uuid",
"positionId": "uuid",
"quantity": "1000000",
"marketValue": "125000000.00"
},
"metadata": {
"correlationId": "uuid",
"causationId": "uuid",
"schemaVersion": "2.1.0"
}
}Version in the event. Schema registry (we use Confluent). Backward compatibility enforced in CI.
Lesson 2: Idempotency Over Exactly-Once
True exactly-once is expensive and often unnecessary. Idempotent consumers are simpler and more robust.
async function handlePositionUpdated(event: PositionUpdatedEvent) {
const idempotencyKey = `position-${event.data.positionId}-${event.eventId}`;
await redis.setnx(idempotencyKey, "processing", { EX: 86400 });
try {
await updateMaterializedView(event.data);
await redis.set(idempotencyKey, "completed");
} catch (error) {
await redis.del(idempotencyKey);
throw error;
}
}Deduplication at the consumer level. Natural idempotency where possible (upserts).
Lesson 3: Materialized Views for Read Patterns
Don't query event streams directly. Build projections.
We maintain 12 materialized views for different access patterns:
- Portfolio summary (real-time)
- Position details (real-time)
- Risk metrics (5-min incremental)
- Regulatory reports (hourly batch)
- Audit log (append-only)
Each view optimized for its query pattern. PostgreSQL for relational, TimescaleDB for time-series, Redis for hot data.
Lesson 4: Observability Is Non-Negotiable
You cannot debug what you cannot see.
Every event flow has:
- Distributed tracing (OpenTelemetry → Jaeger)
- Metrics: lag, throughput, error rate, processing latency
- Alerts: consumer lag > 5min, error rate > 0.1%, processing latency > P99 threshold
- Dead letter queue with automated retry and manual replay
Lesson 5: Schema Evolution Discipline
Breaking changes to event schemas require:
- New event version (v2.2.0)
- Dual-write old + new for 30 days
- Consumer migration
- Deprecate old version
Never delete fields. Mark deprecated. Add new fields as optional.
The Architecture Today
[Producers] → [Kafka Topics] → [Consumer Groups] → [Materialized Views] → [API Layer]
↓
[Schema Registry]
↓
[Dead Letter Queue] → [Replay Tooling]10M events/day. P99 latency < 200ms. 99.99% availability. Zero data loss in 18 months.
Building event-driven systems? We can help.