The Monster in the Basement
Apex Insurance's policy administration system: 2.3M lines of Java, 15 years old, 40 developers afraid to touch it. Deployments took 6 hours. One schema change broke 200 tests.
They didn't need microservices. They needed deployability.
The Strangler Fig Pattern
Don't rewrite. Strangle.
[Legacy Monolith] ←→ [Facade/Proxy] → [New Services]
↓
[Traffic Router]
↓
[Canary Releases]- Identify a bounded context
- Build new service alongside
- Proxy traffic to new service
- Validate, then cut over
- Delete legacy code
- Repeat
Phase 1: The Facade
We built an API gateway (Kong) in front of the monolith. Every request passes through. This gave us:
- Request/response logging
- Rate limiting
- Authentication termination
- Traffic routing rules
Phase 2: Domain Extraction Order
We mapped the domain using Domain-Driven Design. Extraction order:
- Notifications (low risk, high value) - Email, SMS, push
- Document Generation - PDF policies, quotes, letters
- Payment Processing - Already partially externalized
- User Management - Auth, profiles, permissions
- Quoting Engine - Core business logic, high complexity
- Policy Administration - The crown jewel, highest risk
Each extraction: 6-10 weeks. Parallel tracks after phase 1.
Phase 3: Data Synchronization
The hardest part. Shared database = distributed monolith.
Strategies we used:
- Dual Write: Write to both old and new. Reconciliation job nightly.
- Change Data Capture: Debezium streams MySQL binlog → Kafka → new service DB.
- API Facade: New service owns data. Legacy calls new service API.
- Materialized Views: For reporting, build read models from events.
Phase 4: Traffic Migration
Canary releases with feature flags:
# LaunchDarkly flag
newQuotingEngine:
variations:
- legacy: 90%
- new: 10%
targeting:
- internal-users: 100% new
- beta-customers: 50% new
- all: gradual rollout over 4 weeksMetrics gates: error rate, latency, business metrics (quote accuracy, conversion).
Phase 5: Decommission
Legacy code deletion is satisfying but dangerous.
Checklist before delete:
- [ ] Zero traffic for 30 days
- [ ] All tests pass without legacy stubs
- [ ] Documentation updated
- [ ] Team trained on new service
- [ ] Rollback plan tested (it's just a flag flip)
Results After 18 Months
- 42 services extracted
- Deployment frequency: monthly → 50/day
- Lead time: 6 hours → 15 minutes
- Change failure rate: 15% → 0.5%
- MTTR: 4 hours → 12 minutes
- Developer satisfaction: 3.2/10 → 8.7/10
The Playbook
- Start with observability - You can't migrate what you can't measure
- Extract low-risk, high-value first - Build confidence and tooling
- Invest in platform - CI/CD, service mesh, observability, deploy tooling
- Data is the blocker - Plan synchronization before extraction
- Culture eats architecture - Teams must own services end-to-end
Facing a monolith migration? We've done this before.