Adam Burich Open to work

← Portfolio

Featured · Technical write-up

Behavioral danger scoring from overhead highway video

Solo · June–July 2026 · ~25,000 lines of Python across ~125 modules · 286 commits · built in roughly three weeks

An end-to-end system that watches highway traffic from overhead video — drone footage or research datasets — tracks every vehicle, and scores each driver's behavior in real time on a 0–100 Danger Index. The goal is to surface the drivers who are genuinely endangering the people around them (tailgaters, weavers, aggressive cut-ins) rather than crude proxies like raw speed, and to do it in a way that is evaluated against human judgment and explainable signal by signal.

At a glance

The live analyst view: highway footage with every vehicle boxed and colored by its running danger assessment, beside a ranked board of the worst drivers.
The live analyst view mid-run on real drone footage: every vehicle boxed and colored by its running assessment — note the single stable box around the full semi rig — beside the ranked board of the worst drivers among 1,548 tracked, on a feed the system self-calibrated from a cold start in 10 seconds.

The problem

Traffic-safety analysis already uses established surrogate measures such as time-to-collision, time headway, speed, braking, and other kinematic indicators to characterize risky interactions. Individually, however, these measures capture only part of what a human observer considers when judging whether a driver is creating danger for surrounding traffic. This project explores whether overhead video can combine multiple surrogate measures and interaction-level behavioral signals into a continuous, explainable assessment of driving risk.

Overhead video (drones, pole cameras) can see this — every vehicle, every interaction, no instrumentation in the cars. The hard parts are (1) turning pixels into clean, physically-calibrated trajectories, (2) turning trajectories into a defensible danger measure, and (3) testing whether that measure bears a meaningful relationship to human judgments of dangerous driving. This project addresses all three.

What the system does

Two operating modes share one scoring engine:

Offline / research mode ingests the highD dataset — 60 drone recordings of German highways, ~110,000 vehicle trajectories — and replays it through the scorer, driving an interactive visualizer with score-colored bounding boxes, a clickable "hall of shame," and per-driver assessments.

Live mode ingests raw drone video with no ground truth at all. The system detects vehicles, tracks them, calibrates itself (scale, road geometry, traffic direction — no manual setup), converts tracks into metric trajectories in the same format as the research dataset, and scores them in real time behind a web dashboard: live video with score overlays, a ranked driver board, and per-driver behavioral assessments with the evidence for each signal.

The live dashboard: raw feed and annotated analyst view side by side, a ranked board of tracked drivers, and a replay drill-down of one flagged vehicle.
Live mode on a 15-minute drone clip: raw feed (left), analyst view with score-colored boxes (right), the ranked board of 1,737 tracked drivers — and the drill-down replay of one flagged vehicle, rebuilt from retained footage so any driver's episode can be reviewed on demand.

The scoring layer never knows which mode it is in. That is by design: the pipeline is structured so look-ahead is structurally impossible — the scorer's only input is a rolling 3-second buffer, so it behaves identically on recorded and live data, and its outputs are honest about what could have been known at that moment.

video ──▶ tiled detection ──▶ min-cost-flow tracking ──▶ self-calibration ──▶ metric trajectories
                                                                                    │
highD dataset ──────────────────────────────────────────────────────────────────────┤
                                                                                    ▼
                          frame simulator ──▶ rolling 3 s buffer ──▶ signal metrics
                                                                          ▼
                              noisy-OR scorer ──▶ Danger Index + per-driver assessment
                                                                          ▼
                              visualizers · live web dashboard · video export · CSV/NDJSON

01Detection: seeing small vehicles in large frames

Drone frames are 1080p–4K and a car can be a few dozen pixels wide; naive full-frame detection downscales the image and erases exactly the vehicles you need. The detector here slices each frame into overlapping tiles at native resolution, runs a YOLOv8 network pretrained on the VisDrone aerial benchmark on each tile, maps boxes back to frame coordinates, and merges them with class-aware non-maximum suppression, plus one full-frame pass to catch large vehicles straddling tile seams. Everything is resolution-agnostic — the tile grid derives from the actual frame size at runtime.

Two problems needed custom work beyond the pretrained model:

02Tracking as global optimization

Rather than greedy frame-to-frame matching, tracking is formulated as a min-cost network-flow problem: every detection is a node, edges carry motion/appearance/gap costs, and the solver finds the globally optimal set of vehicle paths. Physical structure is encoded in the graph itself — vehicles cannot appear or vanish mid-road for free (a conservation prior with an explicit escape valve), and a velocity-continuity penalty is seeded by mutual-nearest-neighbor optical flow, which stays clean precisely in the ambiguous overtaking cases.

The solve uses Google OR-tools' C++ min-cost-flow solver — ~300× faster than the pure-Python network-simplex baseline (23 s → 0.08 s per clip) with verified-identical optimal cost, with a graceful fallback when OR-tools is absent.

The offline tracker is the reference. For live operation I built a rolling-window variant that re-solves a bounded sub-graph as frames arrive with a 1-second look-ahead, reusing feed-time caches so the per-frame step cost stays flat. Acceptance against the offline reference was explicit and measured: mean score divergence ≤ 8 points and 95th-percentile ≤ 30 on the 0–100 scale, and zero flips in the ranked driver board across the test fleet.

A guiding principle throughout: lose a track cleanly rather than recover it wrongly. A dropped track costs coverage; a wrong recovery silently corrupts every downstream safety metric.

03Self-calibration: no site setup

Scoring needs metric quantities (meters, m/s), but a drone feed arrives as raw pixels with unknown scale and orientation. The live system bootstraps its own geometry from the traffic it observes:

04The scoring engine

The core measure is behavioral, built on surrogate safety measures from the traffic-safety literature — time headway (THW), time-to-collision (TTC), distance headway — plus interaction signals computed over each vehicle's rolling 3-second window: forcing others to brake, cutting in with insufficient gap, weaving, speed vs. surrounding flow (lane-aware), sustained jerk, lateral control variance, and more. Ten signals, each normalized 0–100 and individually explainable.

Design choices that turned out to matter:

Thresholds (Watch at 50, Critical at 80) are set on the index itself — not to an alert budget, a separation I kept deliberate: operational filters (such as top-k review) layer on top of the index rather than distorting it. Whether 80 is the right place for Critical is still an open question; see what's next below.

Danger Index — Watch and Critical thresholds

0 Watch
50
Critical
80
100

05Live operation

The live stack is a working analyst tool, not a demo script:

The dashboard early in a live run: raw feed and annotated analyst view side by side with a delay readout and self-calibration banner, above the running board of drivers.
The dashboard early in a live run: raw feed and annotated analyst view side by side, the true end-to-end delay pinned on screen (+0.6 s behind live), and the self-calibration banner — both carriageways detected, scale resolved to 12.9 px/m — above the running board.

06Validation: initial comparison with human judgment

The current evidence asks a narrower question: does the score's relative ordering broadly agree with human judgments of dangerous driving? This is an initial human-rater evaluation, not an independent validation of the Danger Index.

One trial from the ratings review tool: two drone clips labelled A (score 65.5, the model's pick) and B (31.8), above a table of each anonymized rater's choice, confidence, response time and rationale.
One trial from the real-footage deck in the ratings review tool: clip A (scored 65.5, the model's pick) versus clip B (31.8), with each rater's blind choice, confidence, response time, and free-text rationale. Six of seven raters chose A — concordant with the model.

07Performance engineering

All hot paths were profiled before touching, and every optimization landed only after an equivalence-diff check — outputs compared element-for-element against the pre-change implementation.

ComponentResultVerification
Scoring engine (dense traffic, CPU)~2–4 ms/frame — real-time at 25 fps with ~10× headroomoutput-identical
Track association + smoothing25.9 → 6.9 ms/frame (~3.7×) via cached medians, vectorized smoothing, cached solve-graph assemblybit-identical
Flow solve23 s → 0.08 s per clip (~300×) via OR-tools C++ solveridentical optimal cost
Detection stageFP16 + TensorRT + auto-ROI + tile batching; parallel batch driverbit-identical
Live assessment board62 → 0.35 ms per update at 30k driversoutput-identical
Live vs. offline trackermean Δ ≤ 8 pts, p95 ≤ 30, zero leaderboard flipsmeasured acceptance criteria

Two profiling results went the other way: GPU acceleration of the scorer was evaluated and rejected (the workload is many small windows, not big tensors — the win came from prefix sums and algorithmic complexity instead), and warm-starting the flow solver was refuted by profile (the C++ solve was only ~9% of step time; the assumed bottleneck wasn't the bottleneck).

08Engineering practices

The toolbox GUI: every pipeline stage, visualizer and validation script registered with its parameters and a one-click Run button, with a live log.
The toolbox: every pipeline stage, visualizer, and validation script registered with its parameters and a one-click Run — here mid-launch of the timed live pipeline, the log showing calibration warm-up, TensorRT engines loading, and throughput ramping to ~28 fps.

09Beyond the code

10What's next

A pooled 83.7% from six raters over a two-dozen-pair deck is encouraging evidence, not proof — it measures agreement on ranking, over footage from a handful of sites. The gaps I consider real, and roughly the order I'd attack them:

None of these are hypothetical: they trace to open TODO items, shelved branches with written analyses, and plans already in the repo.

Development status

This is an independently developed research prototype. The implementation, experimental design, scoring decisions, acceptance criteria, and documented limitations are my own. The system remains under active development, particularly around validation and broader-domain testing.

Stack

Python · NumPy · OpenCV · PyTorch · Ultralytics YOLOv8 (VisDrone-pretrained + custom fine-tune) · TensorRT · Google OR-tools · SciPy / pandas / Matplotlib · stdlib HTTP server + vanilla-JS web UIs · highD dataset · BeamNG.drive (simulated drone footage for development) + real drone capture