TL;DR
Braintrust released Nitro, an asynchronous query engine that separates storage I/O from compute scheduling, delivering 2x faster full-text searches across production trace data.
Key Points
- Cold-start queries improved 2-4x; warm queries optimized to match Tantivy baseline performance
- Decouples storage and compute workers via separate concurrency limits, preventing CPU/memory oversubscription
- Aggressive prefetching and request coalescing reduce object-store round trips from 4+ to 1-2 per lookup
- Shingle search optimization eliminates segments before phrase verification, cutting expensive posting list traversals
Why It Matters
As agents query trace history at scale—hundreds of concurrent searches across thousands of traces—query latency directly impacts investigation speed and agent efficiency. Nitro's architectural separation of I/O and compute enables production debugging workflows to handle agent-driven workloads without cache dependency, critical for observability platforms serving AI-heavy deployments.
Source: www.braintrust.dev