Technical overviewData engineering and observability

Sentinel

A streaming telemetry pipeline from host collectors to a day-partitioned archive.

Source code

The repository is private. The architecture, engineering decisions, and verified evidence are documented below.

Role
Solo engineer: collection, streaming, storage, deployment, public interface
Evidence status
Live at sentinel.pesanth.com; collection, streaming, storage, and the dashboard run, with baseline analytics planned
Verified
2026-08-15

01

Purpose and scope

Collects telemetry from two self-hosted machines, streams it through Kafka into a day-partitioned HDFS archive, and reads it back on a public dashboard. Collectors publish every ten seconds, buffer to disk when the broker is down, and replay in order once it returns. It is live at sentinel.pesanth.com.

What it is used for

  • Tell whether a server was actually down, from recorded history.
  • Track processor temperature, memory pressure, and disk headroom over time.
  • Keep a durable record that outlives stream retention.
  • Python
  • Apache Kafka
  • Hadoop HDFS
  • Docker
  • Flask
  • systemd

02

Architecture

Break the pipeline

A simulation of the design, not a live feed. Every interval, threshold and path below is a value the code actually uses.

Simulated clock2026-08-21 23:57:00 UTC40x speed

while broker downreplay in orderservercollectorworkstationcollectorKafka brokerHDFS sinkday partitionsdisk spool

Disk spool

0 readings · est. 0.0 KiB of 64 MiB

Sink batch

0 of 500 records · flush in 60 s of 60

Archive

0 files

Nothing written yet.

Stop the broker before the clock passes 00:00, then start it again.

  • Collector beats are 10 s for system, 30 s for container and 60 s for HTTP checks. The sink flushes at 500 records or 60 seconds, whichever comes first, so at these rates the time trigger always fires and the record cap is never reached.
  • Spool bytes are sized at 680 B per reading, the serialised size of one system reading measured on the workstation collector on 2026-09-06. Container and HTTP readings differ.
  • The real anchor: during an unplanned outage starting 2026-08-21 the spool drained 5.4 MB in about eight seconds. The drain is slowed here so it can be watched.
  • Sentinel's local stack has been stopped since 2026-08-25. sentinel.pesanth.com serves the cached mirror on g7, and nothing is archiving right now.

Collection

Agent

Server collector

systemd service, unprivileged

Agent

Workstation collector

Scheduled task at logon

Durability

Disk spool

Bounded buffer for unsent readings

Publish · keyed by host

Transport

Message broker

Kafka broker

Single node, KRaft mode

Streams

Reading topics

System, container, and endpoint streams

Consume · commit after write

Storage

Consumer

HDFS sink

Commits offsets only after files land

Coordinator

HDFS NameNode

Namespace and block metadata

Archive

Day partitions

Immutable JSON by observation day

Read · WebHDFS

Presentation

Web interface

Dashboard

Host cards, endpoint checks, coverage

Public ingress

Cloudflare Tunnel

Outbound-only ingress, no open port

Text equivalent: Collectors publish readings to Kafka over a private network, buffering to disk when the broker is down; a sink writes newline-delimited JSON into HDFS partitioned by observation day, and a dashboard reads that archive behind an outbound-only Cloudflare Tunnel.

03

Engineering decisions

Partition by observation time

Readings are filed under the day they were taken, not the day they arrived, so a collector that buffered through an outage restores the correct history.

Commit offsets only after the data is durable

The sink writes partition files before committing offsets, so an interruption repeats a batch rather than losing it.

Buffer at the edge

Each collector keeps a bounded local buffer and replays it in order, claiming the file by rename so no reading is dropped.

04

Verification evidence

  • 103 of 104 automated tests pass without Kafka, HDFS, a network, or a container runtime. The one failure is a dashboard test that has failed on main since before 2026-08-23.
  • Edge buffering was verified by interrupting the broker: five readings held on disk, then replayed once it returned.
  • Durability was verified by restarting the platform: the archive grew from nine files to twelve with earlier files intact.
verified-2026-08-15
$ python -m collector --once
5 readings · 0 delivered · 5 failed
$ docker compose start kafka
10 delivered · 5 replayed from disk
✓ /sentinel/raw/system/dt=2026-08-15

05

Limitations