Leonardo Mitsuo Fukuda
Work

SYS/001 · 2026 / Active

Observa

A local distributed-order lab demonstrating diagnosis, compensation, and recovery with end-to-end observability.

Systems designPythonNode.jsKafkaPostgreSQLKubernetesGrafanaEvidence repositoryDocumentation
ORDERKAFKAPAYMENT · INVENTORYOTELGRAFANA

01 / Context & problem

Observa is a local lab for diagnosing and recovering a distributed order journey. Four services exchange domain events; orders, payments, and stock are synthetic. The goal is to make the path from a business state to its trace and related logs visible.

  • Three demonstrable outcomes: CONFIRMED, FAILED, and CANCELLED; the CANCELLED path includes payment compensation.
  • The demo runs on local Kubernetes, with no charging, external notification, or required cloud service.
  • OrderFlow was a read-only functional reference; Observa is a separate project.

02 / Role & scope

The repository presents services, contracts, infrastructure, verification scripts, and runbooks as the project's engineering work. Public evidence supports reviewing behavior and decisions without attributing outcomes to an external commercial operation.

  • Event flow and contracts for Order, Payment, Inventory, and Notification.
  • Outbox, deduplication, and recovery tested with disposable PostgreSQL and in the local cluster.
  • Instrumentation correlating metrics, traces, and logs in Grafana.

03 / System architecture

Order and Notification use Python; Payment and Inventory use Node.js. Kafka transports events, PostgreSQL persists effects, and telemetry reaches an observable stack on local Kubernetes. Processing is at least once; event_id deduplication prevents duplicate effects.

  • Approved payment and reserved stock lead to CONFIRMED; rejected payment leads to FAILED.
  • Unavailable stock after approval requests and confirms a refund before CANCELLED.
  • Normal-path observed ordering is scoped to topic and partition; replay and parking do not preserve the original position.

04 / Data model

Order state and consumer effects are persisted in PostgreSQL. Payment writes its decision, processing marker, and outbox in one transaction. event_id identifies redelivery without repeating the effect.

  • The harness forces a failure between commit and offset to verify redelivery with one effect per event_id.
  • After recreating Inventory, recovery verification requires one reservation and one terminal event.
  • IDs, credentials, and raw evidence are kept under .local/, outside Git; data is synthetic.

05 / Engineering decisions

The lab favors observable failures and reproducible experiments. Outbox and deduplication support at-least-once delivery; the dashboard alone does not prove persisted effects.

  • Transactional PostgreSQL claims in relays were tested with two claimers. ACK before commit can still republish; deduplication remains necessary.
  • KEDA and HPAs are optional. A controlled test attributed Payment scaling to Kafka lag without claiming high availability.
  • Alerts are experimental; one test observed the Payment alert firing and clearing after recovery.

06 / Implementation

The README gives an executable PowerShell path: install tools, start the dedicated minikube profile, check status and contracts, run all three scenarios, and test recovery. The runbook shows how to open Grafana through port-forward.

  • Core commands: ./scripts/mvp.ps1 up, status, contract, demo, and recovery, in that order.
  • check.ps1 covers builds, lint, types, tests, images, audits, secrets, manifests, and Prometheus rules; public CI passed at commit 5d3267e.
  • The observa-spike0 profile is local and dedicated; the README documents prerequisites and removal.

07 / Product experience

The evaluator starts with the demo's final states and continues in the observa-mvp dashboard. A recent exemplar opens a Tempo trace; Related logs shows Loki records; View trace returns to the same execution.

  • The dashboard displays events, errors, duration, and business effects, with exemplars.
  • The runbook supports reproducing status and demo without extra manual steps in the domain flow.
  • Grafana navigation requires visual inspection; the written summary alone does not certify UI clicks.

08 / Results

The public MVP summary records all three completed states. Recovery recorded a CONFIRMED order after Inventory was recreated with one reservation. Post-MVP validation adds relay, scaling, alert, and local failure experiments.

  • A later local baseline completed 30/30 sequential orders; observed p95 values by scenario were 2.309 s, 1.129 s, and 2.228 s.
  • In a clean checkout, the inspection followed an exemplar to a four-service, 21-span trace, 10 logs with the same trace_id, then back to that trace.
  • These are local observations with synthetic data, not production percentiles or an approved SLO.

09 / Retrospective

The project demonstrates diagnosis and recovery under controlled conditions. One node, one broker, and synthetic data limit the conclusion: recreating pods demonstrates eventual recovery, not continuity through an outage or high availability.

  • Replay, parking, and concurrency have documented ordering limits.
  • The experimental hypothesis of 95% of orders within 5 s was not adopted as an SLO.
  • Production claims would require representative environments and loads.