# Dalton AI

> Dalton is the AI Reliability Platform. It joins your code, CI/CD, cloud and observability into one map. It catches failures before your alerts fire and hands you the fix to approve, with the evidence attached. Dalton proposes; your engineers decide what ships.

Dalton is used by engineering, SRE, platform and DevOps teams. It runs on top of the observability and alerting you already have (Datadog, Prometheus, Grafana, PagerDuty and others), read-only by default, and replaces none of it.

## Product

- [Reliability](https://daltonhq.ai/ai-reliability): The product page. Catch production failures before they page you: how Dalton predicts, investigates, proposes and learns. Also defines AI reliability and how it differs from observability and AIOps.
- [Why Dalton](https://daltonhq.ai/why-us): Other tools start at the page; Dalton starts before it. How it compares with dashboards, AIOps and AI SRE agents.
- [Integrations](https://daltonhq.ai/integrations): 51 tools in five kinds (Observability, Kubernetes and cloud, Code and CI/CD, Knowledge and tickets, Chat), read-only by default, no agents or sidecars in your services.
- [Security](https://daltonhq.ai/security): What Dalton can see, keep and touch. SOC 2 Type II, GDPR, PII redacted on the way in, your data never trains a model, and every write needs a person's approval, each time.
- [Customers](https://daltonhq.ai/case-studies): One real incident, start to finish: 57 hours without Dalton, minutes with it.

## Concepts

- [AI SRE](https://daltonhq.ai/ai-sre): What an AI SRE is, how it differs from AIOps and a human SRE, and how to evaluate one. Dalton does this work and starts earlier.
- [System Reliability](https://daltonhq.ai/system-reliability): The outcome: services that stay up, fast and correct while they change.

## What Dalton does

- The problem Dalton exists for is structural, not staffing: you ship hundreds of changes a week, each one moves what normal looks like, and your alert rules only know the normal you wrote down. More people on call won't close that gap.
- It joins your code, CI/CD, cloud and observability into one map, current on every deploy. The map stays current between incidents, so an investigation never starts from zero.
- It sees what is about to break: a slow climb that starts after a deploy is investigated while it is still under every threshold, and failures in places nobody set an alert still surface.
- It works every signal itself, before anyone is paged. A signal is an alert, a status change or a sharp shift in a metric. What reaches a person is the cause, with the evidence attached.
- It hands you the fix as a pull request, with the evidence attached. It does not change production on its own: every integration is read-only by default, and every write needs a person's approval, each time. Access can be revoked whenever you want.
- It learns what your team teaches it, and every lesson records who taught it and when.
- One message per cause, not per symptom, to whoever is on call.
- How to try it: A demo first, then a two-week proof of concept on your own production, read-only on every integration. Walk away at any point; there is nothing to unpick afterwards.

## Writing

- [Blog](https://daltonhq.ai/blog): Essays and engineering notes on AI reliability, production and running agents in real systems.

## Contact

- Book a demo: https://calendar.notion.so/meet/itamarknafo/demo
- Careers: https://daltonhq.ai/careers
