Explore use cases

One workflow. Different operations.

Each example follows the same path: collect the signals you already have, detect what departs from usual behaviour in logs and metrics, then connect the observations into a finding your team can examine.

Use case 02 · Digital services

Checkout slows down, then starts failing.

Customers see slow order pages and failed payments. Payment, inventory and the API gateway all report errors within seconds of each other. Every team sees its own symptom. Which change came first?

Signals you already have.

Logs and metrics are gathered from the systems that produce them, without changing how those systems run.

  • AUTH-SVCAuthentication · Service
  • API-GWAPI gateway · Edge
  • PAYMENT-SVCPayments · Service
  • INVENTORY-SVCInventory · Service
  • DATABASEShared database · Data

Log stream

  1. AUTH-SVCUser login successful userId=8342
  2. API-GWGET /api/orders 200 duration=45ms
  3. PAYMENT-SVCPayment authorized orderId=9832
  4. AUTH-SVCFailed login attempt, repeated 312 times in 60s
  5. API-GWGET /api/orders 500 duration=3200ms
  6. PAYMENT-SVCPayment service timeout
  7. INVENTORY-SVCError connecting to DB
  8. DATABASEConnection pool exhausted (100/100)
  9. API-GWPOST /api/checkout 503 duration=5020ms

Unusual log patterns stand out.

Log models learn the usual events and sequences for each source, then flag lines that depart from that reference.

6 of 9 lines flagged

Log stream with detected anomalies

  1. AUTH-SVCUser login successful userId=8342
  2. API-GWGET /api/orders 200 duration=45ms
  3. PAYMENT-SVCPayment authorized orderId=9832
  4. AUTH-SVCFailed login attempt, repeated 312 times in 60sUnusual volume
  5. API-GWGET /api/orders 500 duration=3200msError and latency
  6. PAYMENT-SVCPayment service timeoutRare event
  7. INVENTORY-SVCError connecting to DBRare event
  8. DATABASEConnection pool exhausted (100/100)Capacity limit
  9. API-GWPOST /api/checkout 503 duration=5020msCustomer impact

Metrics leave their usual range.

Metrics models look across several measurements over time and mark the windows where behaviour changes.

  • Failed logins per minuteAUTH-SVC

    Burst starts first

  • Active DB connectionsDATABASE

    Pool at 100%

  • Payment latency p95PAYMENT-SVC

    Timeouts

  • API 5xx rateAPI-GW

    5xx errors

A lead your team can test.

Anomalies from logs and metrics are placed in order and combined with operational context into one finding.

Context added

  • Service map
  • Deployments
  • Runbooks

Illustrative use case. Records, timings and values are invented to explain the intended workflow. Timing and relationships suggest a lead; they do not prove its cause.

Use case 03 · Cloud platforms

A routine release slows the order API.

A new version of a service rolls out on Kubernetes. Twenty minutes later latency rises and pods restart, but no alert names the release. Is it traffic, the cluster or the change?

Signals you already have.

Logs and metrics are gathered from the systems that produce them, without changing how those systems run.

  • CI/CDDelivery pipeline · Change
  • K8SCluster events · Platform
  • ORDERS-APIOrders service · Service
  • INGRESSIngress gateway · Edge

Log stream

  1. CI/CDRollout started: orders-api v2.14.0 (3 replicas)
  2. K8SPod orders-api-7d9f4-x2kqp Started
  3. ORDERS-APIGET /orders 200 duration=48ms
  4. CI/CDRollout complete: orders-api v2.14.0
  5. ORDERS-APICache client: response buffer grew to 512MB
  6. K8SPod orders-api-7d9f4-x2kqp OOMKilled (limit 1Gi)
  7. K8SBack-off restarting failed container orders-api
  8. INGRESSupstream timed out GET /orders 504

Unusual log patterns stand out.

Log models learn the usual events and sequences for each source, then flag lines that depart from that reference.

4 of 8 lines flagged

Log stream with detected anomalies

  1. CI/CDRollout started: orders-api v2.14.0 (3 replicas)
  2. K8SPod orders-api-7d9f4-x2kqp Started
  3. ORDERS-APIGET /orders 200 duration=48ms
  4. CI/CDRollout complete: orders-api v2.14.0
  5. ORDERS-APICache client: response buffer grew to 512MBNew message type
  6. K8SPod orders-api-7d9f4-x2kqp OOMKilled (limit 1Gi)Rare event
  7. K8SBack-off restarting failed container orders-apiRepeated
  8. INGRESSupstream timed out GET /orders 504Error and latency

Metrics leave their usual range.

Metrics models look across several measurements over time and mark the windows where behaviour changes.

  • Memory per podORDERS-API

    Grows after release

  • Container restartsK8S

    OOM restarts

  • Latency p95INGRESS

    Latency spikes

  • Requests per secondINGRESS

    Traffic unchanged

A lead your team can test.

Anomalies from logs and metrics are placed in order and combined with operational context into one finding.

Context added

  • Deployments
  • Cluster events
  • Release notes

Illustrative use case. Records, timings and values are invented to explain the intended workflow. Timing and relationships suggest a lead; they do not prove its cause.

Your operation

Which service would you start with?

A high-level description is enough: what you operate, the logs and metrics you collect, and where investigation gets difficult. We will discuss whether the approach fits before considering an evaluation.

Discuss your use case