Ongrid

Ongrid

Ongrid is a self-hosted ops AI agent that finds incident root cause from Slack or Telegram using your live telemetry.

68/100MonitorFree planFreemium

If your team lives in Prometheus and Grafana and you're tired of hand-writing queries mid-incident, Ongrid is worth the install. It queries live telemetry and returns an evidence-backed root cause, and the outbound-only tunnel plus read-only audited skills make it defensible for shops that can't ship telemetry to a SaaS vendor. Budget for self-hosting ops — this is not sign-up-and-go.

Verified 40m ago · liveness 68/100 · cite: rightaichoice.com/tools/ongrid

Best for
  • DevOps engineers already running Prometheus, Loki, and Grafana who want to troubleshoot without leaving Slack or Telegram
  • SRE teams cutting MTTR by asking infrastructure questions in natural language and getting evidence back
  • Platform engineering teams wiring custom runbooks and audited read-only skills into chat workflows
  • Security-conscious orgs that require on-prem or air-gapped deployment with an outbound-only tunnel and full audit trail
Not ideal for
  • Teams that want a fully managed SaaS with zero self-hosting overhead
  • Shops with no existing Prometheus, Loki, or Grafana stack — setup cost lands before any value does
  • Anyone looking for a general-purpose AI assistant outside of ops and incident work
Visit Website

AdvancedThe one-liner install targets roughly 5 minutes on Linux amd64 or ARM64 from the v0.16.0 release, and the first-boot checklist lets you point at existing Prometheus, Loki, and Tempo endpoints instead of the bundled ones. Realistically, budget an afternoon for channel wiring in Slack or Telegram and skill/permission review, and longer if you need air-gapped or on-prem scope.CLI · WebAPI availableVerified 40m ago
Pricing
Free plan
FreemiumFree tier3 plans6 hidden costs
Learning curve
Advanced
The one-liner install targets roughly 5 minutes on Linux amd64 or ARM64 from the v0.16.0 release, and the first-boot checklist lets you point at existing Prometheus, Loki, and Tempo endpoints instead of the bundled ones. Realistically, budget an afternoon for channel wiring in Slack or Telegram and skill/permission review, and longer if you need air-gapped or on-prem scope.
Runs on
CLIWeb
API available · 11 integrations
Who it's for
On-call SRE at a mid-size SaaS companyPlatform engineer maintaining a self-hosted observability stackSecurity-conscious infra lead at a regulated company
Live sentiment
Is Ongrid actually worth it?

We scan live Reddit threads, YouTube comments, X posts, G2 reviews and other communities — and hand you an honest verdict in under a minute.

  • Honest verdict, not marketing
  • Real pros & cons from real users
  • Attributed quotes with receipts
Run a free scan

3 free scans · no card needed

Skip it if

Skip Ongrid if you have no existing Prometheus/Loki/Tempo stack and no one to own Docker Compose upgrades, the MySQL instance, and the outbound tunnel — the self-hosting work lands before any incident gets faster.

The 30-second take
Biggest gripe

Model calls are on you: Ongrid is bring-your-own LLM, so your Anthropic, OpenAI, Gemini, or relay bill is billed separately and scales with how much the agent queries.

Price reality

The Free tier (1,000 AI calls, 1 integration) suits a solo SRE or a single-service pilot. Pro is per-user contact-sales pricing and fits a platform team running multiple environments and regions. Enterprise is a custom quote for air-gapped or on-prem scope. Against hosted AIOps copilots you avoid per-GB telemetry ingest fees but pay in engineering time; against a plain Grafana-only setup the tool itself is effectively free while your LLM provider bill does the real spending.

In short

Ongrid — Ongrid is a self-hosted ops AI agent that finds incident root cause from Slack or Telegram using your live telemetry. Best for DevOps engineers already running Prometheus, Loki, and Grafana who want to troubleshoot without leaving Slack or Telegram, SRE teams cutting MTTR by asking infrastructure questions in natural language and getting evidence back, Platform engineering teams wiring custom runbooks and audited read-only skills into chat workflows. Free to use.

What's new in Ongrid

Checked 9 days ago

Across the latest 1 update: 1 feature update.

What people actually say about Ongrid — is it worth it?

We ran a structured research pass across product reviews, community discussions, and post-purchase forum threads to surface the patterns vendors won't publish themselves. Below: the recurring strengths, the hidden costs people mention most, and the cohort that consistently regrets adopting this tool.

9 mentions across 3 sources (Hacker News, GitHub, Lemmy) · researched Jul 3, 2026.

12% positive88% critical

Average across the 3 sources that answered — each source counts once, not each post.

Recurring strengths
  • +Chat-native troubleshooting keeps teams in Slack/Telegram without context switching.
  • +Infrastructure-aware root cause analysis using live data from multiple cloud providers.
  • +Automated remediation via runbooks and scripts reduces manual intervention.
  • +Supports major cloud providers: AWS, Azure, GCP.
  • +Integrates with popular monitoring tools like Datadog, New Relic, PagerDuty.
Recurring frustrations
  • −Core features like host_bash tool often fail to execute commands.
  • −Configuration UI has bugs where input fields disappear after ten seconds.
  • −Permission errors after adding nodes prevent monitoring from loading.
  • −Deployment in air-gapped environments fails due to Docker timeout.
  • −Lack of multi-replica support creates a performance bottleneck.
Patterns worth knowing
Reliability and bugginess in basic operations are the top community concern.
Seen on GitHub
Users want more integrations, especially with databases and high-availability setups.
Seen on GitHub
Deployment complexity and network dependency frustrate self-hosters.
Seen on GitHub
Learning curve
intermediateProductive in ~A few hours
Hidden costs people mention
  • • Custom action definitions and approval workflows may require paid plan; no transparent pricing published

Viability Score

68/100
Monitor

How well maintained and how widely used is Ongrid? Built from what the vendor actually publishes (docs, changelog, tutorials, integrations, pricing), whether the site is live, and how much real users discuss it. How we calculate this

Recent activity
90
Traction
90
Site health
95
User sentiment
12
What the vendor publishes
40

Last calculated: October 2026

How we score →

Key Features

  • Root cause analysis across metrics, logs, traces, topology, and source code
  • Generates and executes PromQL, LogQL, and TraceQL queries from chat
  • Chat-based incident troubleshooting in Slack, Telegram, Larksuite, DingTalk, and WeCom
  • Two-way channels with per-channel locale configuration
  • Webhook channel support for custom event sources
  • Read-only host probing via bash and host_probe_* skills
  • Around 26 audited skills plus custom runbooks
  • Expand topology graphs to add context to an answer
  • Evidence-backed answers grounded in live telemetry, not chat transcripts
  • Bring your own LLM: Anthropic, OpenAI, GLM, DeepSeek, Gemini, Kimi, or any OpenAI-compatible relay
  • One-liner install for Linux amd64/ARM64 (Ubuntu 22.04+, Debian 12+, RHEL 9, Rocky 9)
  • Self-hosted Docker Compose deployment bundling manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant
  • Outbound-only geminio tunnel per host: zero open inbound ports, no jumpbox
  • Full audit trail for every agent action
  • Multi-environment and multi-region support

About Ongrid

FreemiumAdvancedAPI availableCLI · Web

Ongrid is an ops AI agent for teams that already run their own observability stack. It plugs into Prometheus, Grafana, Loki, Tempo, OpenTelemetry, and Qdrant, then reasons over metrics, logs, traces, topology, source code, and your knowledge base to answer incident questions in plain language — returning a grounded, evidence-backed answer rather than a chat transcript. It lives where on-call engineers already are: Slack, Telegram, Larksuite, DingTalk, or WeCom, with per-channel locale settings. The agent calls tools as needed — bash, host_probe_*, query_promql, expand_topology among roughly 26 audited read-only skills — and generates PromQL, LogQL, and TraceQL on the fly, expanding topology graphs to add context. Every call is logged, so an audit trail exists without extra configuration. The deployment model is the real dividing line. Each host dials out through a single outbound geminio tunnel: no inbound ports, no jumpbox. Install is a one-liner on Linux amd64/ARM64 (Ubuntu 22.04+, Debian 12+, RHEL 9, Rocky 9), or `docker compose up` for the full stack of manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant. You bring the model — Anthropic, OpenAI, GLM, DeepSeek, Gemini, Kimi, or any OpenAI-compatible relay. That puts Ongrid in a different bracket from hosted AIOps copilots that keep your telemetry in their cloud. It assumes you already have an observability stack and are willing to run it yourself; if you don't, the setup cost is real and lands before any value does.

Behind the Verdict

We'd reach for Ongrid when the observability stack is already in place and the bottleneck is response time, not instrumentation. It fits an SRE team that wants to ask "what changed in the last 30 minutes?" from a phone and get the actual PromQL and trace span behind the answer, not a summary of the thread. The read-only skill model is the part that earns trust. Around 26 audited skills, every call logged — you can hand this to an on-call rotation without handing it the keys to production. Custom runbooks extend it for team-specific procedures. Where it bites: self-hosting is not a footnote. You're running the manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant, or you're pointing at your existing stack and doing the wiring. That's a platform team's afternoon, not a solo dev's lunch break. Pick Ongrid over a managed AIOps copilot when data residency or procurement rules mean telemetry cannot leave your network. Pick the managed option when you want someone else to operate the agent and you don't mind the data leaving. Know what it isn't: Ongrid doesn't do on-call scheduling, paging, or alert routing. It sits alongside your PagerDuty or Opsgenie, not on top of them.

Researching Ongrid? Get your full AI stack in 60 seconds.

Free, no signup — tell us your goal and get tools matched to your budget & existing stack.

Real-world workflow fit

Concrete scenarios for the personas Ongrid actually fits — and what changes day-one when you adopt it.

On-call SRE at a mid-size SaaS company

Paged at 2am, you open Slack and ask the Ongrid channel why p99 latency doubled on the checkout service. The agent generates PromQL against your live Prometheus, pulls correlated Loki logs, expands the topology graph to find the upstream dependency that changed, and replies with the evidence.

Outcome: You identify the cause from your phone in one thread instead of opening five browser tabs, and every query the agent ran is logged for the morning review.

Platform engineer maintaining a self-hosted observability stack

You run `docker compose up` to bring up the manager, MySQL, Prometheus, Loki, Tempo, Grafana, and Qdrant, then point each host at the outbound geminio tunnel and wire Slack as the first channel.

Outcome: The stack comes up from one command with no inbound ports opened, and you switch to your existing Prom/Loki/Tempo endpoints from the first-boot checklist.

Security-conscious infra lead at a regulated company

You deploy Ongrid with read-only audited skills only — bash, host_probe_*, query_promql, expand_topology — and keep telemetry on your own network through the outbound-only tunnel.

Outcome: Engineers get natural-language infrastructure answers while every agent action lands in an audit trail and no telemetry leaves your environment.

Use Cases

  • Ask 'What caused the spike in latency?' in Slack and get a root cause backed by PromQL and topology evidence.
  • Investigate a misconfigured AWS security group by querying live state from chat.
  • Expand the service topology graph mid-incident to see what changed upstream.
  • Run an audited read-only host probe with bash and host_probe_* skills during triage.
  • Keep an audit trail of every query and action the agent ran for post-incident review.

Models Under the Hood

AnthropicOpenAIGLMDeepSeekGeminiKimi

as of 2026-10-03

Limitations

  • Ongrid requires a self-hosted deployment via Docker Compose and supports specific Linux distributions (Ubuntu 22.04+, Debian 12+, RHEL 9, Rocky 9).
  • It depends on observability stacks you already run (Prometheus, Loki, Tempo, OpenTelemetry, Qdrant) and on your chat channel providers; answer quality tracks the quality of the monitoring data behind it.
  • Because model access is bring-your-own, you pay the LLM provider separately from Ongrid itself.
  • Someone on your team must own upgrades, the MySQL instance, the containers, and the outbound geminio tunnel.
  • Ongrid does not handle on-call scheduling, paging, or alert routing — it is a troubleshooting layer beside your incident management stack, not a replacement for it.

as of 2026-09-15

Verification history

We have re-verified Ongrid 8 times since . Each pass re-reads the vendor's own pages and re-checks every listed field against that evidence; passes where nothing had changed are marked as such.

  1. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  2. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  3. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  4. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  5. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it
  6. — re-verified summary, description, our verdict, our analysis, pricing model, pricing tiers, features, integrations, who it suits, who should skip it

Showing the 6 most recent of 8 verification passes.

Free to cite with attribution — this page re-verifies continuously.

12-month cost

Project the real annual outlay, including the implied monthly cost when only an annual tier is published.

Annual total
Free
Over 12 months
Effective monthly
Free
Billed monthly

Vendor list price only. Add-on usage, seat overages, and contract minimums are surfaced under Hidden costs & gotchas.

Plans compared

For each published Ongrid tier: who it actually fits, and what it adds vs. the previous tier. Cross-reference the cost calculator above for projected annual outlay.

Free

$0/mo

Ideal for

A solo SRE or single-service pilot that wants to trial root-cause analysis from Slack or Telegram without a procurement cycle.

What this tier adds

Starting tier: 1,000 AI calls, 1 integration, read-only audited skills, and bring-your-own LLM.

Pro

Contact sales

Ideal for

A platform team running several environments or regions that needs shared custom runbooks across channels.

What this tier adds

Adds expanded multi-environment and multi-region support, custom runbooks, per-channel locale configuration, and webhook channels.

Enterprise

Custom

Ideal for

Regulated or security-constrained orgs that need air-gapped or on-prem deployment and custom contractual terms.

What this tier adds

Adds custom terms and deployment scope, air-gapped or on-prem support, the self-hosted Docker Compose stack, and full audit trail coverage.

Hidden costs & gotchas

What the public pricing page doesn't put in bold. Captured from pricing-page footnotes, contract terms, and recurring complaints.

  • Model calls are on you: Ongrid is bring-your-own LLM, so your Anthropic, OpenAI, Gemini, or relay bill is billed separately and scales with how much the agent queries.
  • Free covers only 1,000 AI calls and 1 integration, so a busy on-call rotation burns through the cap and pushes you toward Pro faster than the tier list suggests.
  • Pro is priced per user, so the cost climbs with every engineer you add to the chat workspace rather than with actual incident volume.
  • You absorb the infrastructure bill for the self-hosted stack — MySQL, the manager, Prometheus, Loki, Tempo, Grafana, and Qdrant all run on your own compute.
  • Enterprise features such as air-gapped or on-prem deployment support and custom terms sit behind a custom quote, so compliance-driven buyers cannot stay on the published tiers.
  • Because you run the containers, upgrades and tunnel maintenance are unbudgeted engineering hours that a managed SaaS would have absorbed.

Where the pricing makes sense

The company stage and team size where Ongrid's pricing actually pencils out — and where peers do it cheaper.

The Free tier (1,000 AI calls, 1 integration) suits a solo SRE or a single-service pilot. Pro is per-user contact-sales pricing and fits a platform team running multiple environments and regions. Enterprise is a custom quote for air-gapped or on-prem scope. Against hosted AIOps copilots you avoid per-GB telemetry ingest fees but pay in engineering time; against a plain Grafana-only setup the tool itself is effectively free while your LLM provider bill does the real spending.

Setup time & first value

How long it actually takes to get something useful out of Ongrid — broken out by persona, not the marketing-page minute.

The one-liner install targets roughly 5 minutes on Linux amd64 or ARM64 from the v0.16.0 release, and the first-boot checklist lets you point at existing Prometheus, Loki, and Tempo endpoints instead of the bundled ones. Realistically, budget an afternoon for channel wiring in Slack or Telegram and skill/permission review, and longer if you need air-gapped or on-prem scope.

Switching to or from Ongrid

How to bring data in from common predecessors and how to get it back out — written for the switcher, not the buyer.

Migrating in
  • →From manual Grafana and Prometheus querying: install via the one-liner, then switch to your existing Prom/Loki/Tempo from the first-boot checklist.
  • →From a hosted AIOps copilot: deploy the Docker Compose stack on your own hosts and wire Slack or Telegram as the channel.
  • →From an ad-hoc runbook wiki: port the steps into Ongrid custom runbooks and let the audited read-only skills execute them.
Migrating out
  • ↗To a managed AIOps platform: export audit logs and channel history, then re-point your alerting and chat workflows at the hosted service.
  • ↗To plain Grafana alerting: keep your existing Prometheus, Loki, and Tempo backends and drop the chat agent layer.
  • ↗To a general-purpose chat assistant: replace the audited skills with manual queries, accepting the loss of the logged audit trail.

Integrations

SlackTelegramLarksuiteDingTalkWeComPrometheusGrafanaLokiTempoOpenTelemetryQdrant

Resources & Guides

Tutorials & Learning

YouTube returned 6 videos for “Ongrid”, and we withheld 6: 6 could not be judged, because “Ongrid” is a single word that other videos use for other things. We are showing none, because we could not prove any of them are about Ongrid.

Official links

Tools that pair well with Ongrid

Common stack mates teams adopt alongside Ongrid, with the specific reason each pairing earns its keep.

Featured Head-to-Head Comparisons

Alternatives to Ongrid

View all
Corelayer

Corelayer

AI-native incident response that finds production root causes, cuts alert noise, and opens fix PRs — deployable on-prem or in your cloud.

Contact SalesTry
Deeptrace

Deeptrace

AI SRE agent that investigates production alerts and posts evidence-backed root causes in Slack within minutes.

FreemiumTry
Interfere

Interfere

Interfere is AI production monitoring that finds bugs, explains root causes, and suggests fixes before users report them.

Contact SalesTry

Frequently Asked Questions

Used Ongrid? Help shape our editorial sentiment research.