Ongrid
Ongrid is a self-hosted ops AI agent that finds incident root cause from Slack or Telegram using your live telemetry.
If your team lives in Prometheus and Grafana and you're tired of hand-writing queries mid-incident, Ongrid is worth the install. It queries live telemetry and returns an evidence-backed root cause, and the outbound-only tunnel plus read-only audited skills make it defensible for shops that can't ship telemetry to a SaaS vendor. Budget for self-hosting ops — this is not sign-up-and-go.
Verified 40m ago · liveness 68/100 · cite: rightaichoice.com/tools/ongrid
- DevOps engineers already running Prometheus, Loki, and Grafana who want to troubleshoot without leaving Slack or Telegram
- SRE teams cutting MTTR by asking infrastructure questions in natural language and getting evidence back
- Platform engineering teams wiring custom runbooks and audited read-only skills into chat workflows
- Security-conscious orgs that require on-prem or air-gapped deployment with an outbound-only tunnel and full audit trail
- Teams that want a fully managed SaaS with zero self-hosting overhead
- Shops with no existing Prometheus, Loki, or Grafana stack — setup cost lands before any value does
- Anyone looking for a general-purpose AI assistant outside of ops and incident work
We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.
- Honest verdict, not marketing
- Real pros & cons from real users
- Attributed quotes with receipts
3 free scans · no card needed
Skip Ongrid if you have no existing Prometheus/Loki/Tempo stack and no one to own Docker Compose upgrades, the MySQL instance, and the outbound tunnel — the self-hosting work lands before any incident gets faster.
Model calls are on you: Ongrid is bring-your-own LLM, so your Anthropic, OpenAI, Gemini, or relay bill is billed separately and scales with how much the agent queries.
The Free tier (1,000 AI calls, 1 integration) suits a solo SRE or a single-service pilot. Pro is per-user contact-sales pricing and fits a platform team running multiple environments and regions. Enterprise is a custom quote for air-gapped or on-prem scope. Against hosted AIOps copilots you avoid per-GB telemetry ingest fees but pay in engineering time; against a plain Grafana-only setup the tool itself is effectively free while your LLM provider bill does the real spending.
In short
Ongrid — Ongrid is a self-hosted ops AI agent that finds incident root cause from Slack or Telegram using your live telemetry. Best for DevOps engineers already running Prometheus, Loki, and Grafana who want to troubleshoot without leaving Slack or Telegram, SRE teams cutting MTTR by asking infrastructure questions in natural language and getting evidence back, Platform engineering teams wiring custom runbooks and audited read-only skills into chat workflows. Free to use.
What's new in Ongrid
Checked 9 days agoAcross the latest 1 update: 1 feature update.
What people actually say about Ongrid — is it worth it?
We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.
9 mentions across 3 sources (Hacker News, GitHub, Lemmy) · researched Jul 3, 2026.
Average across the 3 sources that answered — each source counts once, not each post.
- +Chat-native troubleshooting keeps teams in Slack/Telegram without context switching.
- +Infrastructure-aware root cause analysis using live data from multiple cloud providers.
- +Automated remediation via runbooks and scripts reduces manual intervention.
- +Supports major cloud providers: AWS, Azure, GCP.
- +Integrates with popular monitoring tools like Datadog, New Relic, PagerDuty.
- −Core features like host_bash tool often fail to execute commands.
- −Configuration UI has bugs where input fields disappear after ten seconds.
- −Permission errors after adding nodes prevent monitoring from loading.
- −Deployment in air-gapped environments fails due to Docker timeout.
- −Lack of multi-replica support creates a performance bottleneck.
- • Custom action definitions and approval workflows may require paid plan; no transparent pricing published
Viability Score
How well maintained and how widely used is Ongrid? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this
Last calculated: October 2026
How we score →Key Features
- Root cause analysis across metrics, logs, traces, topology, and source code
- Generates and executes PromQL, LogQL, and TraceQL queries from chat
- Chat-based incident troubleshooting in Slack, Telegram, Larksuite, DingTalk, and WeCom
- Two-way channels with per-channel locale configuration
- Webhook channel support for custom event sources
- Read-only host probing via bash and host_probe_* skills
- Around 26 audited skills plus custom runbooks
- Expand topology graphs to add context to an answer
- Evidence-backed answers grounded in live telemetry, not chat transcripts
- Bring your own LLM: Anthropic, OpenAI, GLM, DeepSeek, Gemini, Kimi, or any OpenAI-compatible relay
- One-liner install for Linux amd64/ARM64 (Ubuntu 22.04+, Debian 12+, RHEL 9, Rocky 9)
- Self-hosted Docker Compose deployment bundling manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant
- Outbound-only geminio tunnel per host: zero open inbound ports, no jumpbox
- Full audit trail for every agent action
- Multi-environment and multi-region support
About Ongrid
Ongrid is an ops AI agent for teams that already run their own observability stack. It plugs into Prometheus, Grafana, Loki, Tempo, OpenTelemetry, and Qdrant, then reasons over metrics, logs, traces, topology, source code, and your knowledge base to answer incident questions in plain language — returning a grounded, evidence-backed answer rather than a chat transcript. It lives where on-call engineers already are: Slack, Telegram, Larksuite, DingTalk, or WeCom, with per-channel locale settings. The agent calls tools as needed — bash, host_probe_*, query_promql, expand_topology among roughly 26 audited read-only skills — and generates PromQL, LogQL, and TraceQL on the fly, expanding topology graphs to add context. Every call is logged, so an audit trail exists without extra configuration. The deployment model is the real dividing line. Each host dials out through a single outbound geminio tunnel: no inbound ports, no jumpbox. Install is a one-liner on Linux amd64/ARM64 (Ubuntu 22.04+, Debian 12+, RHEL 9, Rocky 9), or `docker compose up` for the full stack of manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant. You bring the model — Anthropic, OpenAI, GLM, DeepSeek, Gemini, Kimi, or any OpenAI-compatible relay. That puts Ongrid in a different bracket from hosted AIOps copilots that keep your telemetry in their cloud. It assumes you already have an observability stack and are willing to run it yourself; if you don't, the setup cost is real and lands before any value does.
Behind the Verdict
We'd reach for Ongrid when the observability stack is already in place and the bottleneck is response time, not instrumentation. It fits an SRE team that wants to ask "what changed in the last 30 minutes?" from a phone and get the actual PromQL and trace span behind the answer, not a summary of the thread. The read-only skill model is the part that earns trust. Around 26 audited skills, every call logged — you can hand this to an on-call rotation without handing it the keys to production. Custom runbooks extend it for team-specific procedures. Where it bites: self-hosting is not a footnote. You're running the manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant, or you're pointing at your existing stack and doing the wiring. That's a platform team's afternoon, not a solo dev's lunch break. Pick Ongrid over a managed AIOps copilot when data residency or procurement rules mean telemetry cannot leave your network. Pick the managed option when you want someone else to operate the agent and you don't mind the data leaving. Know what it isn't: Ongrid doesn't do on-call scheduling, paging, or alert routing. It sits alongside your PagerDuty or Opsgenie, not on top of them.
Researching Ongrid? Get your full AI stack in 60 seconds.
Free, no signup — tell us your goal and get tools matched to your budget & existing stack.
Real-world workflow fit
Concrete scenarios for the personas Ongrid actually fits — and what changes day-one when you adopt it.
Paged at 2am, you open Slack and ask the Ongrid channel why p99 latency doubled on the checkout service. The agent generates PromQL against your live Prometheus, pulls correlated Loki logs, expands the topology graph to find the upstream dependency that changed, and replies with the evidence.
Outcome: You identify the cause from your phone in one thread instead of opening five browser tabs, and every query the agent ran is logged for the morning review.
You run `docker compose up` to bring up the manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant, then point each host at the outbound geminio tunnel and wire Slack as the first channel.
Outcome: The stack comes up from one command with no inbound ports opened, and you switch to your existing Prom/Loki/Tempo endpoints from the first-boot checklist.
You deploy Ongrid with read-only audited skills only — bash, host_probe_*, query_promql, expand_topology — and keep telemetry on your own network through the outbound-only tunnel.
Outcome: Engineers get natural-language infrastructure answers while every agent action lands in an audit trail and no telemetry leaves your environment.
Use Cases
- Ask 'What caused the spike in latency?' in Slack and get a root cause backed by PromQL and topology evidence.
- Investigate a misconfigured AWS security group by querying live state from chat.
- Expand the service topology graph mid-incident to see what changed upstream.
- Run an audited read-only host probe with bash and host_probe_* skills during triage.
- Keep an audit trail of every query and action the agent ran for post-incident review.
Models Under the Hood
as of 2026-10-03
Limitations
- Ongrid requires a self-hosted deployment via Docker Compose and supports specific Linux distributions (Ubuntu 22.04+, Debian 12+, RHEL 9, Rocky 9).
- It depends on observability stacks you already run (Prometheus, Loki, Tempo, OpenTelemetry, Qdrant) and on your chat channel providers; answer quality tracks the quality of the monitoring data behind it.
- Because model access is bring-your-own, you pay the LLM provider separately from Ongrid itself.
- Someone on your team must own upgrades, the MySQL instance, the containers, and the outbound geminio tunnel.
- Ongrid does not handle on-call scheduling, paging, or alert routing — it is a troubleshooting layer beside your incident management stack, not a replacement for it.
as of 2026-09-15
Verification history
We have re-verified Ongrid 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
- — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
Showing the 6 most recent of 8 verification passes.
Free to cite with attribution — this page re-verifies continuously.
12-month cost
Project the real annual outlay, including the implied monthly cost when only an annual tier is published.
Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.
Plans compared
For each published Ongrid tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.
Free
$0/mo
Ideal for
A solo SRE or single-service pilot that wants to trial root-cause analysis from Slack or Telegram without a procurement cycle.
What this tier adds
Starting tier: 1,000 AI calls, 1 integration, read-only audited skills, and bring-your-own LLM.
Pro
Contact sales
Ideal for
A platform team running several environments or regions that needs shared custom runbooks across channels.
What this tier adds
Adds expanded multi-environment and multi-region support, custom runbooks, per-channel locale configuration, and webhook channels.
Enterprise
Custom
Ideal for
Regulated or security-constrained orgs that need air-gapped or on-prem deployment and custom contractual terms.
What this tier adds
Adds custom terms and deployment scope, air-gapped or on-prem support, the self-hosted Docker Compose stack, and full audit trail coverage.
Where the pricing makes sense
The company stage and team size where Ongrid's pricing actually pencils out — and where peers do it cheaper.
The Free tier (1,000 AI calls, 1 integration) suits a solo SRE or a single-service pilot. Pro is per-user contact-sales pricing and fits a platform team running multiple environments and regions. Enterprise is a custom quote for air-gapped or on-prem scope. Against hosted AIOps copilots you avoid per-GB telemetry ingest fees but pay in engineering time; against a plain Grafana-only setup the tool itself is effectively free while your LLM provider bill does the real spending.
Setup time & first value
How long it actually takes to get something useful out of Ongrid — broken out by persona, not the marketing-page minute.
The one-liner install targets roughly 5 minutes on Linux amd64 or ARM64 from the v0.16.0 release, and the first-boot checklist lets you point at existing Prometheus, Loki, and Tempo endpoints instead of the bundled ones. Realistically, budget an afternoon for channel wiring in Slack or Telegram and skill/permission review, and longer if you need air-gapped or on-prem scope.
Switching to or from Ongrid
How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.
- →From manual Grafana and Prometheus querying: install via the one-liner, then switch to your existing Prom/Loki/Tempo from the first-boot checklist.
- →From a hosted AIOps copilot: deploy the Docker Compose stack on your own hosts and wire Slack or Telegram as the channel.
- →From an ad-hoc runbook wiki: port the steps into Ongrid custom runbooks and let the audited read-only skills execute them.
- ↗To a managed AIOps platform: export audit logs and channel history, then re-point your alerting and chat workflows at the hosted service.
- ↗To plain Grafana alerting: keep your existing Prometheus, Loki, and Tempo backends and drop the chat agent layer.
- ↗To a general-purpose chat assistant: replace the audited skills with manual queries, accepting the loss of the logged audit trail.
Integrations
Resources & Guides
Tutorials & Learning
YouTube returned 6 videos for “Ongrid”, and we withheld 6: 6 could not be judged, because “Ongrid” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Ongrid.
Official links
Tools that pair well with Ongrid
Common stack mates teams adopt alongside Ongrid, with the specific reason each pairing earns its keep.
Corelayer
AI-native incident response that finds production root causes, cuts alert noise, and opens fix PRs — deployable on-prem or in your cloud.
Deeptrace
AI SRE agent that investigates production alerts and posts evidence-backed root causes in Slack within minutes.
Interfere
Interfere is AI production monitoring that finds bugs, explains root causes, and suggests fixes before users report them.
Featured Head-to-Head Comparisons
Ongrid vs Spider Cloud
Ongrid and Spider Cloud serve completely different purposes: Ongrid is an ops AI agent for troubleshooting infrastructure from chat, while Spider Cloud is a web scraping API for AI agents. Choose Ongrid if you need self-hosted incident response with query generation and remediation; choose Spider Cloud if you need cost-effective, reliable web data for RAG or LLMs.
Ongrid vs Presto Voice
Presto Voice and Ongrid serve completely different domains: Presto Voice is purpose-built for QSR drive-thru automation with proven revenue lift, while Ongrid is an ops AI agent for DevOps teams to troubleshoot infrastructure from chat. Choose based on your business vertical—restaurant ops or IT operations—as there is no overlap. Both require contacting sales for pricing.
Ongrid vs Temporal Ai
Choose Temporal AI if you need reliable, stateful orchestration for complex AI agents or microservices with automatic retries and recovery, especially if you prefer an open-source platform with a managed cloud option. Choose Ongrid if your primary need is infrastructure-level troubleshooting from chat and you already run a Prometheus/Loki/Grafana stack on-premise. They overlap minimally.
Alternatives to Ongrid
View allFrequently Asked Questions
Categories
Best-of guides
Used Ongrid? Help shape our editorial sentiment research.