Hi everyone,
I'm currently running an Apache Druid 37.0.0 cluster with 20+ nodes and am looking to improve our monitoring and operational workflows using AI.
I'm interested in learning how others in the community are approaching AI-assisted monitoring and operations for Druid. Specifically:
Monitoring
Which Druid and infrastructure metrics do you consider the most critical to monitor?
If you're using Prometheus, which metrics and alerting rules have proven the most valuable in production?
Are there any dashboards (Grafana or otherwise) that you would recommend as a starting point?
AI / LLM Integration
Has anyone integrated Druid monitoring with AI agents (OpenAI, Claude, Gemini, etc.) for anomaly detection, root-cause analysis, or operational assistance?
Have you found this useful in production, and if so, what architecture or tooling are you using?
MCP (Model Context Protocol)
Are there any mature MCP servers or integrations that work well for infrastructure operations?
I'm particularly interested in capabilities such as:
Executing controlled operational actions (for example, restarting Druid services after approval)
Reading Prometheus metrics
Querying Grafana dashboards
Inspecting logs
Diagnosing Druid issues
If you've built something similar, I'd appreciate hearing about your experience and any lessons learned.
I'm looking for practical production experiences, recommended tools, architectures, or open-source projects that have worked well.
Thanks in advance!