Point quality in tracking: ghost marks, thresholds, trimming, smoothing¶
For a short, plain explanation of all the tracking parameters (what to measure, how to set them, what changes) see Tracking parameters; this page goes deeper on the ghost marks.
This page explains the quality marks the two_phase tracker uses by default, how to
change the two thresholds, when to change them, and how to test a choice on
your own data. Evidence and measurements are in
docs/plans/2026-09-30-tracker-plan.md (sections 1 and 4, A10) and bench/.
1. What problem this solves¶
About 5–10% of the 3D points that camera matching produces are ghosts: made from
blobs of different particles. A ghost often links into a long, smooth, wrong
trajectory. On the synthetic benchmark ghosts cause about 60% of the remaining
velocity error of two_phase and nearly all of its wrong links.
A ghost's camera rays usually meet worse than a real particle's rays. The distance by which they miss each other is the ray miss distance (rcm). It is the best ghost signal we found (AUC 0.88 per trajectory). No single cut-off works, so the tracker uses it softly, as a ghost probability between 0 and 1 for every point.
2. How the mark is computed¶
- rcm of every stored 3D point (
openptv2.point_quality.frame_rcm), from the targets and calibrations in the run store. - rcm grows away from the volume centre on real data (about 0.035 mm near the centre,
0.058 mm at 50–60 mm), so it is divided by a straight line in the distance from the
centre, fitted per run on the 4-camera points (
fit_scale). - The relative rcm and the number of cameras give the probability
(
ghost_probability; a 3-camera point is far more likely to be a ghost than a 4-camera point with the same relative rcm). - The probability is stored in the run store (
quality/frame_NNNNNN,RunStore.read_point_quality) so later steps can use it without recomputing.
Model q_model (track section; default rcm_blob, alternative rcm):
rcm: steps 1-3 above.rcm_blob: also uses the brightness agreement of the blobs across cameras. A real particle shows a similar brightness in all cameras that see it (spread of the log brightness about 0.15 on real data); a ghost combines blobs of different particles (about 0.4). The brightness of each camera is first divided by that camera's typical value (its run median), because real cameras differ in gain by 0.13-0.17 in log, as much as the spread itself. It is a logistic model on [relative rcm, 3-camera flag, log spread], fitted on synthetic data whose per-camera brightness scatter matches real data (synth_bench build --amp-jitter 0.18). If the blobs carry no brightness (constant values) the plugin falls back torcmwith a message.- Measured on four independent realistic synthetic cases (against
rcm): velocity error −0.005 (sparse), −0.015 (×2 density), −0.039 (×4 density), −0.001 (clustered with frame skip ×4); ghost fraction 1 to 2 points lower; true points kept at most 0.007 lower. On real wp2: 4.6% fewer trajectories, mean length 20.2 → 20.8, suspect trajectories 10.4% → 9.7%, jump share unchanged. The share of points confidently flagged (probability above 0.5) is about 5% in both the synthetic test and wp2.
Nothing is deleted at this stage.
3. The two thresholds¶
Set them in the track section (YAML / GUI parameters) of the two_phase tracker.
| parameter | default | what it does | switch off |
|---|---|---|---|
q_seed |
0.2 | A point whose ghost probability is above this may not start a trajectory (an existing trajectory can still pick it up). It is also the limit for q_young. |
q_seed: null |
q_young |
3 | A trajectory with fewer than this many points may not continue onto a point above q_seed. Established trajectories are not affected. |
q_young: 0 |
q_model |
rcm_blob |
ghost model: rcm_blob (rcm + camera count + blob brightness agreement) or rcm |
q_model: rcm if your blobs have no usable brightness or the cameras are saturated |
blob_gate |
none | brightness continuity gate: a candidate whose blob brightness (mean over the cameras that see both points) changed by more than this in log units is not a candidate. A real particle keeps its brightness in each camera from frame to frame (median change 0.06, p90 0.21 on CompleteTest wp2; a random neighbour 0.33). | off. Use 0.5 for frame-skipped or low-frame-rate data (velocity error −0.007 at skip ×4 and ×8, −0.015 clustered; at most 0.005 fewer points kept; it cuts 0.3% of real links on wp2); 0.3 gains more and loses more points; 0.8 does little. Neutral on normal data (−0.001 to −0.003), so not a default |
q_weight |
0 | Multiplies the link cost by 1 + q_weight * probability. Measured: no effect (links onto doubtful points are rarely contested). Leave at 0. |
0 |
To switch the whole feature off: q_seed: null and q_young: 0. It also switches off
by itself, with a message, when the experiment has no calibrations (the mark needs the
camera models).
The values 0.2 / 3 were chosen from 16 combinations on five synthetic cases (normal, real-jitter, clustered with frame skip ×4, 4× density, frame skip ×8). They are the strongest setting where no case lost more than 0.007 of the true points kept. Mean change against switching it off: velocity error −0.017, true points kept −0.004, ghost fraction −0.021. Checked on four real recordings (CompleteTest wp1 and wp2, lv_multi wp4 and wp5): 5–21% fewer trajectories, 3–10% longer mean length, fewer long trajectories that no frame sees with all four cameras, jump share unchanged or lower.
4. When to change them¶
The choice is a trade-off: a lower q_seed and a higher q_young remove more
ghosts and lower the error, and also remove a few real points.
| mean change vs off (5 synthetic cases) | q_young 3 |
6 | 10 |
|---|---|---|---|
q_seed 0.10 |
err −0.026, pts −0.012 | −0.029, −0.019 | −0.032, −0.026 |
q_seed 0.15 |
−0.022, −0.007 | −0.025, −0.011 | −0.025, −0.015 |
q_seed 0.20 (default) |
−0.017, −0.004 | −0.020, −0.007 | −0.021, −0.009 |
q_seed 0.30 |
−0.011, −0.002 | −0.013, −0.003 | −0.013, −0.004 |
| your data | change |
|---|---|
| Many 3-camera points (above ~45%), dense seeding, many short suspect trajectories; accuracy matters more than the number of points | Accuracy mode: q_seed: 0.15, q_young: 6. Clustered flow gains most from a larger q_young (up to 10). |
| Few ghosts: good calibration, mostly 4-camera points (3-camera share below ~20%), low density | Raise q_seed to 0.3, or switch it off. The gain is small and points are not worth losing. |
| Every real point counts (e.g. you compute statistics that need all samples) | q_seed: 0.3, q_young: 3. |
| Very short recordings (below ~30 frames) | The scale line needs at least 200 four-camera points; with fewer it falls back to one global value. Check the flagged share (section 5). |
| Poor or drifting calibration | rcm is large everywhere; the per-run scale compensates, but if the flagged share (section 5) is far above 20% the table is mismatched: raise q_seed or switch off. |
| Very fast flow / large frame skip | No change needed: tested up to skip ×8 (error −0.008). |
Always combine with the two-hop link confirmation (on by default). Its tolerance is
set from the data when the data are sparse and noise-dominated (8 × the median kink,
about 0.5–0.6 mm/frame at the real jitter level) and is the fixed 0.3 otherwise; see
docs/two-phase-tracking.md. On the benchmark the confirmation has two parts: the
dead-end rule (a track may not end by stepping onto a stranger) is worth 0.02–0.1 of
velocity error everywhere, while the kink test must follow the noise: a fixed 0.3 cuts
true links as soon as the jitter exceeds the real level (2× jitter: 0.3072 with 0.3,
0.2372 with the automatic value). On lv_multi wp5 the
jump steps (velocity change above 0.3 mm in depth) were 8.4% without it and 0.0%
with confirm_tol: 0.3, confirm_ends: true; the quality marks add to that, they
do not replace it.
5. How to test a threshold on your data¶
Sanity check, no truth needed. Count the flagged share from the stored marks:
import numpy as np
from openptv2.storage import RunStore
s = RunStore.open("res/run.zarr", mode="r")
g = np.concatenate([s.read_point_quality(f) for f in range(first, last + 1)])
print((g > 0.2).mean(), (g > 0.3).mean(), (g > 0.5).mean())
Expected on well-behaved data: about 10% above 0.2, 4–6% above 0.3, 0.6–2% above 0.5 (synthetic 10.3 / 5.7 / 2.0%; CompleteTest wp1/wp2 9.9 / 4.4 / 0.6%; lv_multi 18 / 10 / 2.2% because half of its points have only three cameras). A flagged share far above this means the data differ from what the table was built on.
Answer-free scores on real data (scripts/realism_metrics.py), before and after a
change of the thresholds. Keep a setting only if:
jump_sharedoes not rise (it should fall);never4_share(long trajectories never seen by all four cameras, the best real-data hint of ghost trajectories) falls;- the number of trajectories falls and
mean_lenrises (fewer fragments); jitter_z_p90does not rise.
uv run python scripts/real_retrack.py base 300 q_seed=null q_young=0 # scratch copy
uv run python scripts/real_retrack.py new 300 q_seed=0.15 q_young=6
uv run python scripts/realism_metrics.py /private/tmp/claude-501/real_runs/new/res/run.zarr --last 300
With known truth (synthetic cases from scripts/synth_bench.py): sweep a grid and
read velocity error against true points kept, tuning on the real-jitter case and
checking on the others; never tune on real data.
uv run python scripts/synth_bench.py build --level L1 --sigma-px 0.08
uv run python scripts/synth_bench.py sweep --cases CASE --trackers "two_phase+q_seed=0.15+q_young=6" -j 4 --no-eval
uv run python scripts/synth_bench.py eval --case CASE --trackers two_phase "two_phase+q_seed=0.15+q_young=6"
uv run python scripts/tune_summary.py q_seed q_young
6. After tracking: trimming doubtful ends and the smoother¶
- Trim.
openptv2.quality_post.trim_doubtful_ends(trajid, frame, ghost, thr=0.3, max_trim=3, max_frac=0.25)removes up to three end points of a trajectory whose ghost probability is abovethr(never more than a quarter of the trajectory per end). On the synthetic cases it lowers the velocity error by another 0.014–0.04 for 0.010 fewer points kept atthr0.3 (0.2: more error gain, 0.018 fewer points). It is optional and costs points; lowerthrfor more accuracy, raise it to keep more. Short trajectories (lv_multi: about 8 frames) lose 9–16% of their points at 0.3/0.2: use 0.5 there or skip it. - Smoother.
openptv2.quality_post.weighted_savgolis a Savitzky–Golay fit at the true frame times (the flowtracks filter assumes consecutive frames), with all points of a short trajectory fitted when it is shorter than the window (dropping one point to get an odd window made the error 0.008 worse), and optional point weights. With unit weights it lowers the velocity error by 0.002–0.008 on every case. Quality weights(1−g)^pand spike removal gave nothing (<0.001), so use unit weights. - openptv-cloud uses both in its
poststep; see its configuration reference (trajectories.smoothing_method,trajectories.trim_doubtful).
7. Limits¶
- The probability table was measured on synthetic data (about 8% ghosts); real data may have a different ghost rate. The real-data checks above are the guard.
- The synthetic ghosts are almost stationary (static false targets); real data have few stationary trajectories. Do not build a ghost rule on speed.
- It cannot remove ghosts that have good rcm; about 40% of the possible ghost gain is reached (all ghosts removed: velocity error 0.1465 vs 0.1845 with the marks, 0.1977 without, on the real-jitter case).