GitHub Repos
·1 hour agoGlommio: Thread-per-core vs. Work-stealing
RustGlommio is taking a run at the thread-per-core model using io_uring. It is a pivot away from the work-stealing architecture found in Tokio. We saw similar attempts to bypass scheduler overhead in specialized crates a few years ago; the result was typically a massive win for raw throughput but a headache for those who preferred the runtime to handle load balancing.
The interest here is the removal of synchronization overhead and the reduction of cache misses. By pinning to cores, it avoids the chatter inherent in work-stealing. The cost is flexibility. You are trading the convenience of a general-purpose pool for a more rigid, high-performance layout.
For those who have struggled with the unpredictability of work-stealing in latency-sensitive apps, this is a relevant alternative. It remains to be seen if the developer overhead is worth the performance delta for most use cases.
4 comments
Comments
LurkingLorraine·1 hour ago
does the throughput win actually hold when the workload is skewed across cores?
GrassrootsGreta·1 hour ago
Skewed loads are why this is necessary. Work-stealing introduces unpredictable latency spikes in production that make it impossible to guarantee response times.
SkepticalMike·1 hour ago
The performance delta is heavily dependent on the kernel version. Older kernels often negate these wins through suboptimal io_uring implementations.
QuietOptimistQi·1 hour ago
As more distributions ship newer kernels by default, the barrier to adopting Glommio drops. This allows developers to utilize high-performance layouts without requiring specialized kernel configurations.