Agent
Server collector
systemd service, unprivileged
Technical overviewData engineering and observability
A streaming telemetry pipeline from host collectors to a day-partitioned archive.
Source code
The repository is private. The architecture, engineering decisions, and verified evidence are documented below.
01
Collects telemetry from two self-hosted machines, streams it through Kafka into a day-partitioned HDFS archive, and reads it back on a public dashboard. Collectors publish every ten seconds, buffer to disk when the broker is down, and replay in order once it returns. It is live at sentinel.pesanth.com.
02
A simulation of the design, not a live feed. Every interval, threshold and path below is a value the code actually uses.
Simulated clock2026-08-21 23:57:00 UTC40x speed
Disk spool
0 readings · est. 0.0 KiB of 64 MiBSink batch
0 of 500 records · flush in 60 s of 60Archive
0 filesNothing written yet.
Stop the broker before the clock passes 00:00, then start it again.
Agent
systemd service, unprivileged
Agent
Scheduled task at logon
Durability
Bounded buffer for unsent readings
Message broker
Single node, KRaft mode
Streams
System, container, and endpoint streams
Consumer
Commits offsets only after files land
Coordinator
Namespace and block metadata
Archive
Immutable JSON by observation day
Web interface
Host cards, endpoint checks, coverage
Public ingress
Outbound-only ingress, no open port
03
Readings are filed under the day they were taken, not the day they arrived, so a collector that buffered through an outage restores the correct history.
The sink writes partition files before committing offsets, so an interruption repeats a batch rather than losing it.
Each collector keeps a bounded local buffer and replays it in order, claiming the file by rename so no reading is dropped.
04
$ python -m collector --once
5 readings · 0 delivered · 5 failed
$ docker compose start kafka
10 delivered · 5 replayed from disk
✓ /sentinel/raw/system/dt=2026-08-1505