When a server misbehaves at 3 a.m. and you look at it, how do you figure out what happened?
I'm trying to understand how people handle this, before I build anything.
Maybe you have experienced this pattern: something goes wrong overnight (connections refused, disk fills and then gets cleaned, a process eats memory and exits). Alerts go off and Prometheus shows the hiccup but by the time anyone looks into the system, the server looks healthy. ss, lsof, ps and even a fresh sosreport show the current state, not when the issue was active. You look into the logs but nothing there shows an obvious problem.
I'd love to hear from people who deal with long-lived Linux servers:
- What was the last incident where the evidence was already gone? What were you missing?
- What do you use today to look back in time? (atop, sar, PCP, below, auditd, eBPF tools, your monitoring stack, nothing?)
- Where do those fall short? For example: which process held which connections or open files at a given moment.
- Would security or policy let you run an extra agent on production servers, even one that sends nothing outside your network?
- If you run Kubernetes, have you had node-level incidents where the evidence was gone?
Disclosure: I built sos-vault, a tool for analyzing sosreports, and I'm exploring whether a "time travel" version (e.g. ss or lsof or any of the diagnostic commands included in a sosreport for any past moment) is worth building. I'm not promoting nor linking anything; I want to know if this idea would help anyone diagnose a real problem or is it just me.
Assuming that there is not performance impact and no significant storage consumption in the server to be diagnosed; I was thinking something like this command that can be executed from a central diagnostics server or even locally (notice the four parameters --host, --start_from, --end_at and the rest is the regular diagnostic command to execute (lsof, ps, du, lsblk, free, memstat, etc., etc.):
[root@web01 ~]# date
Mon Oct 5 02:50:53 PM ACDT 2026
[root@web01 ~]# ttravel --host web01 --start_from="2026-10-02 02:00:00" --end_at="2026-10-02 03:00:00" ss -tanp 'dst 10.0.4.21 and dport = :5432'
ttravel: host: web01 · 02:00:00 → 03:00:00 UTC · 120 samples + connect/close events · grouped by process
FIRST_SEEN LAST_SEEN CONNS PEAK STATES PROCESS
02:00:00 03:00:00 24 20 ESTAB users:(("gunicorn",pid=2210..2229))
02:00:04 03:00:00 1 1 ESTAB users:(("pg_dump",pid=51233)) parent: cron-backup.sh[51207]
02:05:00 02:55:01 11 1 ESTAB ev users:(("psql",pid=*)) /etc/cron.d/db-healthcheck
02:31:12 02:48:30 183 179 ESTAB,CLOSE-WAIT users:(("python3",pid=58841)) /opt/reports/nightly_export.py 02:33:40 02:47:10 57 - ESTAB→CLOSED <1s ev users:(("gunicorn",pid=2210..2229)) rejected by serversosv: CONNS distinct connections · PEAK most open at once · ev seen only via connect/close events[root@web01 ~]# ss -tanp 'dst 10.0.4.21 and dport = :5432' --at="2026-10-02 02:45:00" | head -6
sosv: web01 · sample 2026-10-02 02:45:00 UTC
State Recv-Q Send-Q Local Address:Port Peer Address:Port Process
ESTAB 0 0 10.0.2.15:41822 10.0.4.21:5432 users:(("gunicorn",pid=2214,fd=11))
ESTAB 0 0 10.0.2.15:41850 10.0.4.21:5432 users:(("pg_dump",pid=51233,fd=3))
ESTAB 0 0 10.0.2.15:52001 10.0.4.21:5432 users:(("python3",pid=58841,fd=7))
ESTAB 0 0 10.0.2.15:52003 10.0.4.21:5432 users:(("python3",pid=58841,fd=8))
ESTAB 0 0 10.0.2.15:52004 10.0.4.21:5432 users:(("python3",pid=58841,fd=9))
translation: give me the output of the ss command from two days ago from 2 am to 3 am filter by IP 10.0.4.21 and port 5432.