Model job status with timeouts, run IDs and ordered events
A job can finish while its last visible status still says running. It can also stop sending updates while the display keeps showing a reassuring success. The problem is that a single label often mixes two separate facts: the last result we received and how recently we received a valid event.
This tutorial builds a dependency-free Python simulator that keeps those facts separate. It is a generic software example from FutureDevice, a desk-hardware retailer. It does not connect to, test or demonstrate FutureDevice hardware, and it does not provide an agent integration. You can run it without buying anything.
Define what the display actually means
The model records a last-known phase: unknown, running, waiting, succeeded or failed. It also records local time since the latest accepted event. When that age reaches the timeout, the display becomes DISCONNECTED while retaining the last-known phase in a separate field. Here DISCONNECTED is a conservative display label meaning no recently accepted event; it is not a network diagnosis. A quiet producer and a broken connection look identical to this model.
A heartbeat updates receipt freshness without changing the phase. If the last phase was waiting, a heartbeat still displays WAITING. If no phase has arrived, a heartbeat displays UNKNOWN. It must not invent running, completion or progress.
Choose a small, explicit event contract
Each event has a run_id, a positive integer seq and a kind. The producer owns one increasing sequence for all events in a run, including heartbeats. The consumer explicitly selects a run with begin(); an event from an unfamiliar run cannot select itself. Use a fresh run ID for every new execution, including retries that represent a new job.
Duplicate and lower-sequence events are ignored, so a delayed running event cannot overwrite a newer waiting event. Gaps are allowed. This is not a reordering buffer: if sequence 3 arrives before sequence 2, the latter will be discarded. Use an ordered stream or repeated authoritative snapshots in a real application where every transition matters. A heartbeat can overtake an important phase event on an unordered transport; this example will not reconstruct the missing phase.
For one selected run, the first accepted terminal result wins. After succeeded or failed, subsequent phase events are ignored even if their sequence is higher. Heartbeats remain acceptable. This policy handles terminal-versus-running precedence, but it does not reconcile contradictory results or choose the result with the greatest producer sequence. A correction requires a separate policy; a real retry should get a new run ID.
Save the complete simulator
Use Python 3.8 or later. Save the following as status_state.py. It uses only the Python standard library. No package install, credentials or network access is needed.
#!/usr/bin/env python3
"""Generic software status simulator; no hardware interfaces or dependencies."""
import argparse
from dataclasses import dataclass
import json
import math
import sys
import time
KINDS = frozenset({'running', 'waiting', 'succeeded', 'failed', 'heartbeat'})
TERMINAL = frozenset({'succeeded', 'failed'})
@dataclass(frozen=True)
class Event:
run_id: str
seq: int
kind: str
def __post_init__(self):
if not isinstance(self.run_id, str) or not self.run_id.strip():
raise ValueError('run_id must be a nonempty string')
if type(self.seq) is not int or self.seq < 1:
raise ValueError('seq must be a positive integer')
if not isinstance(self.kind, str) or self.kind not in KINDS:
raise ValueError('unknown event kind')
class Status:
def __init__(self, timeout=10, clock=time.monotonic):
if not math.isfinite(timeout) or timeout <= 0:
raise ValueError('timeout must be finite and positive')
self.timeout, self.clock = timeout, clock
self.run_id, self.phase, self.seq, self.last_seen = None, 'unknown', 0, None
self.used_runs = set()
def begin(self, run_id):
"""Explicit local selection; an event cannot select a different run."""
if not isinstance(run_id, str) or not run_id.strip():
raise ValueError('run_id must be a nonempty string')
if run_id in self.used_runs:
raise ValueError('use a fresh run_id for every new run')
self.used_runs.add(run_id)
self.run_id, self.phase, self.seq, self.last_seen = run_id, 'unknown', 0, None
def accept(self, event):
if event.run_id != self.run_id:
return 'ignored: different run'
if event.seq <= self.seq:
return 'ignored: duplicate or out of order'
# The first terminal result wins. Only heartbeat may follow it.
if self.phase in TERMINAL and event.kind != 'heartbeat':
return 'ignored: terminal result is fixed'
self.seq, self.last_seen = event.seq, self.clock()
if event.kind != 'heartbeat':
self.phase = event.kind
return 'accepted'
def snapshot(self):
age = None if self.last_seen is None else self.clock() - self.last_seen
disconnected = age is None or age >= self.timeout
return {
'run_id': self.run_id,
'last_known_phase': self.phase,
'connection': 'disconnected' if disconnected else 'recent_event',
'display': 'DISCONNECTED' if disconnected else self.phase.upper(),
'seq': self.seq,
'age_seconds': None if age is None else round(age, 3),
}
def demo():
now = [0.0]
state = Status(timeout=5, clock=lambda: now[0])
def show(label):
print(json.dumps({'step': label, **state.snapshot()}))
state.begin('demo-001')
show('selected; no source event yet')
state.accept(Event('demo-001', 1, 'running'))
show('running event')
now[0] = 5.0
show('timeout boundary')
state.accept(Event('demo-001', 2, 'heartbeat'))
show('heartbeat restores freshness; phase still running')
state.accept(Event('demo-001', 3, 'waiting'))
show('waiting for user')
state.accept(Event('demo-001', 4, 'failed'))
show('failed result')
state.accept(Event('demo-001', 5, 'running'))
show('late running cannot undo failure')
state.begin('demo-002')
state.accept(Event('demo-001', 6, 'succeeded'))
show('new run ignores old run result')
state.accept(Event('demo-002', 1, 'succeeded'))
show('new run succeeds')
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--demo', action='store_true')
parser.add_argument('--timeout', type=float, default=10)
args = parser.parse_args()
if args.demo:
demo()
return
state = Status(timeout=args.timeout)
for line in sys.stdin:
if not line.strip():
continue
try:
command = json.loads(line)
if not isinstance(command, dict):
raise ValueError('command must be a JSON object')
op = command.get('op')
if op == 'begin':
state.begin(command['run_id'])
result = 'selected'
elif op == 'event':
result = state.accept(Event(command['run_id'], command['seq'], command['kind']))
elif op == 'view':
result = 'snapshot'
else:
raise ValueError('op must be begin, event or view')
print(json.dumps({'result': result, **state.snapshot()}), flush=True)
except (ValueError, KeyError, TypeError) as error:
print(json.dumps({'error': str(error)}), flush=True)
if __name__ == '__main__':
main()
Run the deterministic demo
python3 status_state.py --demo
The demo advances an injected clock instead of sleeping. Its displays are DISCONNECTED, RUNNING, DISCONNECTED, RUNNING, WAITING, FAILED, FAILED, DISCONNECTED and SUCCEEDED. The fourth step is receipt-freshness recovery after a heartbeat, not proof that a job progressed. The seventh step shows a later running event failing to undo a terminal failure. Selecting the next run then clears the previous phase and sequence.
Try your own event stream
Run the command below, then enter one JSON object per line. After beginning a run, send an event, wait longer than five seconds and enter a view command to observe the timeout. The CLI prints a snapshot only when input arrives; it does not redraw the screen on a timer. Press Control-D to end input on macOS or Linux.
python3 status_state.py --timeout 5
{"op":"begin","run_id":"build-001"}
{"op":"event","run_id":"build-001","seq":1,"kind":"running"}
{"op":"view"}
{"op":"event","run_id":"build-001","seq":2,"kind":"heartbeat"}
{"op":"event","run_id":"build-001","seq":3,"kind":"failed"}
{"op":"event","run_id":"build-001","seq":4,"kind":"running"}
The last event returns an ignored result, and the phase stays failed. Ignored events do not advance the stored sequence or refresh the receipt timestamp. This prevents a stream of duplicates, wrong-run events or rejected transitions from indefinitely hiding a timeout.
Test failure paths, not just the happy path
Save the following beside the simulator as test_status_state.py and run python3 -m unittest -v. The tests use an injected clock to check the exact timeout boundary without waiting. They also run the actual CLI in a subprocess, ensuring malformed input produces an error while the next valid command still works.
import json
from pathlib import Path
import subprocess
import sys
import unittest
from status_state import Event, Status
class StatusTests(unittest.TestCase):
def setUp(self):
self.now = 0.0
self.state = Status(5, lambda: self.now)
self.state.begin('a')
def event(self, seq, kind, run='a'):
return self.state.accept(Event(run, seq, kind))
def test_selection_is_not_evidence_of_connection(self):
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
self.assertEqual(self.state.phase, 'unknown')
def test_exact_timeout_boundary_and_phase_retention(self):
self.event(1, 'waiting')
self.now = 4.999
self.assertEqual(self.state.snapshot()['display'], 'WAITING')
self.now = 5
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
self.assertEqual(self.state.phase, 'waiting')
def test_duplicate_and_out_of_order_do_not_refresh(self):
self.event(2, 'running')
self.now = 4
self.event(2, 'heartbeat')
self.event(1, 'failed')
self.now = 5
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
self.assertEqual(self.state.phase, 'running')
self.assertEqual(self.state.seq, 2)
def test_heartbeat_recovers_without_fabricating_progress(self):
self.event(1, 'waiting')
self.now = 8
self.event(2, 'heartbeat')
self.assertEqual(self.state.snapshot()['display'], 'WAITING')
self.assertEqual(self.state.snapshot()['age_seconds'], 0)
def test_heartbeat_before_phase_keeps_unknown(self):
self.event(1, 'heartbeat')
self.assertEqual(self.state.snapshot()['display'], 'UNKNOWN')
def test_first_terminal_wins_both_directions(self):
for first, second in [('failed', 'succeeded'), ('succeeded', 'failed')]:
with self.subTest(first=first):
state = Status(5, lambda: self.now)
state.begin(first)
state.accept(Event(first, 1, first))
for seq, kind in enumerate(['running', 'waiting', second], 2):
self.assertIn('terminal', state.accept(Event(first, seq, kind)))
self.assertEqual(state.phase, first)
self.assertEqual(state.seq, 1)
def test_stale_terminal_retained_and_heartbeat_recovers(self):
self.event(1, 'failed')
self.now = 7
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
self.event(2, 'heartbeat')
self.assertEqual(self.state.snapshot()['display'], 'FAILED')
def test_terminal_rejection_does_not_extend_freshness(self):
self.event(1, 'succeeded')
self.now = 4
self.event(2, 'running')
self.now = 5
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
def test_explicit_new_run_ignores_old_events_and_resets_sequence(self):
self.event(50, 'failed')
self.state.begin('b')
self.event(99, 'succeeded')
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
self.assertEqual(self.state.seq, 0)
self.event(1, 'running', 'b')
self.assertEqual(self.state.snapshot()['display'], 'RUNNING')
def test_wrong_run_cannot_select_itself_or_refresh(self):
self.event(1, 'running')
self.now = 5
self.event(100, 'heartbeat', 'b')
self.assertEqual(self.state.snapshot()['display'], 'DISCONNECTED')
self.assertEqual(self.state.run_id, 'a')
def test_run_id_reuse_rejected_even_after_another_run(self):
self.state.begin('b')
with self.assertRaises(ValueError):
self.state.begin('a')
self.assertEqual(self.state.run_id, 'b')
def test_invalid_inputs(self):
for seq in [0, -1, True, 1.5, '2']:
with self.subTest(seq=seq), self.assertRaises(ValueError):
Event('a', seq, 'running')
for timeout in [0, -1, float('inf'), float('nan')]:
with self.subTest(timeout=timeout), self.assertRaises(ValueError):
Status(timeout)
with self.assertRaises(ValueError):
Event('', 1, 'running')
for kind in ['green', []]:
with self.assertRaises(ValueError):
Event('a', 1, kind)
def test_cli_survives_bad_command_and_outputs_state(self):
commands = '\n'.join(['not json', json.dumps({'op': 'begin', 'run_id': 'cli'}),
json.dumps({'op': 'event', 'run_id': 'cli', 'seq': 1, 'kind': 'waiting'}),
json.dumps({'op': 'view'})])
result = subprocess.run([sys.executable, str(Path(__file__).with_name('status_state.py'))],
input=commands, text=True, capture_output=True, check=True)
output = [json.loads(line) for line in result.stdout.splitlines()]
self.assertIn('error', output[0])
self.assertEqual(output[-1]['display'], 'WAITING')
self.assertEqual(len(output), 4)
if __name__ == '__main__':
unittest.main()
The assertions cover run selection without source evidence, stale duplicate events, heartbeat recovery without phase changes, both directions of contradictory terminal results, terminal freshness after a timeout, old-run events after a retry and run-ID reuse within the process. These are the boundaries that tend to produce misleading displays.
Understand what receipt freshness cannot establish
The clock is time.monotonic(), so a system wall-clock adjustment does not move the elapsed receipt time backwards. However, a higher-sequence event delayed in a queue looks fresh when it arrives. There is no producer timestamp, queue-age check, authenticated source or end-to-end liveness handshake. A recent heartbeat proves only that this process recently accepted that event. It does not prove that the job is healthy, its phase is current or a device is connected.
The selected run, sequence, terminal result and set of used run IDs all live in memory. Restarting the process forgets them. The used-ID set also grows for the lifetime of the process; this compact demo is not a long-lived multi-tenant service. Durable checkpoints, bounded retention and replay protection need an explicit design before production use.
The code is single-threaded and has no transport, authentication, source-age filtering, hardware adapter or persistent storage. It assumes a monotonic injected clock and one authoritative sequence per run. It accepts gaps rather than recovering lost transitions. If a producer restarts its sequence counter, select a genuinely new run ID instead of silently accepting a counter reset.
Keep the integration boundary visible
To connect a real system, first decide which component owns run identity and event order. Then document how it reports waiting, terminal outcomes and missed updates. Test reconnects, delayed delivery and a producer restart before letting a display influence decisions. Present the phase and freshness together; color alone should not be the only way to distinguish them.
A physical indicator would require its own verified interface and adapter. This tutorial supplies neither. Its useful output is a small, inspectable state contract that can be tested independently of any desk accessory or agent framework.