TL;DR
Engineering team optimized 15.8M inference tasks by eliminating HTTP boundaries, fixing multiprocessing deadlocks, and implementing idempotent distributed processing.
Key Points
- Per-task latency reduced from 45 seconds to 17 seconds (62% improvement)
- Total runtime: 2 months → 6 days by inlining classification service as library
- Solved fork() deadlock with native ML libraries by switching to spawn() method
- At-least-once SQS delivery at 500 workers required idempotency checks at storage layer, not task level
Why It Matters
This post demonstrates critical patterns for scaling ML inference at population scale: eliminating network boundaries when you control both sides, understanding multiprocessing hazards with native libraries, and implementing idempotency correctly in distributed systems. These lessons apply directly to any team running large-scale batch ML workloads on Kubernetes or similar platforms.
Source: engineering.prod.whoop.com