AI-powered Monitoring & Agentic Operations for Druid 37.0.0

18 views
Skip to first unread message

Unknown

unread,
Jul 29, 2026, 1:14:23 AM (yesterday) Jul 29
to Druid User

Hi everyone,

I'm currently running an Apache Druid 37.0.0 cluster with 20+ nodes and am looking to improve our monitoring and operational workflows using AI.

I'm interested in learning how others in the community are approaching AI-assisted monitoring and operations for Druid. Specifically:

  1. Monitoring

    • Which Druid and infrastructure metrics do you consider the most critical to monitor?

    • If you're using Prometheus, which metrics and alerting rules have proven the most valuable in production?

    • Are there any dashboards (Grafana or otherwise) that you would recommend as a starting point?

  2. AI / LLM Integration

    • Has anyone integrated Druid monitoring with AI agents (OpenAI, Claude, Gemini, etc.) for anomaly detection, root-cause analysis, or operational assistance?

    • Have you found this useful in production, and if so, what architecture or tooling are you using?

  3. MCP (Model Context Protocol)

    • Are there any mature MCP servers or integrations that work well for infrastructure operations?

    • I'm particularly interested in capabilities such as:

      • Executing controlled operational actions (for example, restarting Druid services after approval)

      • Reading Prometheus metrics

      • Querying Grafana dashboards

      • Inspecting logs

      • Diagnosing Druid issues

    • If you've built something similar, I'd appreciate hearing about your experience and any lessons learned.

I'm looking for practical production experiences, recommended tools, architectures, or open-source projects that have worked well.

Thanks in advance!

Reply all
Reply to author
Forward
0 new messages