Our billing pipeline was suddenly slow. The culprit was a hidden bottleneck in ClickHouse
A partitioning change in a petabyte-scale ClickHouse cluster led to stalled billing jobs due to severe lock contention in the query planner, which was resolved by identifying the issue and creating upstream patches.
MAIN POINTS
- Partitioning change caused critical billing jobs to stall in ClickHouse cluster.
- Standard metrics failed to reveal any obvious errors initially.
- Severe lock contention was identified in ClickHouse's query planner.
- Upstream patches were developed to address and fix the issue.
TAKEAWAYS
- Monitoring tools may not always detect underlying issues in complex systems.
- Lock contention can significantly impact performance in database systems.
- Identifying root causes requires deep investigation beyond surface-level metrics.
- Developing and applying patches can effectively resolve critical system issues.