Add benchmark Makefile for eval and Codabench submission. - #436
Draft
AlexBodner wants to merge 60 commits into
Draft
Add benchmark Makefile for eval and Codabench submission.#436AlexBodner wants to merge 60 commits into
AlexBodner wants to merge 60 commits into
Conversation
Introduces benchmark/ with make targets for setup, tune, eval, submit, and upload-codabench on MOT17, SportsMOT, and DanceTrack. Submit uses submit_yolox.py with library defaults; eval uses tracker_flags.py for per-tracker CLI parameters. Co-authored-by: Cursor <cursoragent@cursor.com>
…com/roboflow/trackers into feat/benchmark-codabench-submission
…com/roboflow/trackers into feat/benchmark-codabench-submission
Contributor
There was a problem hiding this comment.
Pull request overview
Adds a repo-local benchmarking workflow intended to reproduce/refresh the tracker comparison numbers by orchestrating data preparation, tuning, tracking, evaluation, and (where applicable) Codabench submissions. It also updates the docs comparison page to reflect updated detection sources (notably for DanceTrack).
Changes:
- Introduce a new
benchmark/directory with a Makefile-driven pipeline and helper scripts for MOT-format prep, tracking, formatting, uploads, and score aggregation. - Add Codabench submission + polling tooling (pure stdlib HTTP client) and MOT17-specific submission formatting.
- Update
docs/trackers/comparison.mdwording about which datasets use YOLOX vs oracle detections.
Reviewed changes
Copilot reviewed 12 out of 12 changed files in this pull request and generated 9 comments.
Show a summary per file
| File | Description |
|---|---|
| docs/trackers/comparison.md | Updates benchmark/detection-source notes and DanceTrack detection wording. |
| benchmark/Makefile | Orchestrates setup, prep, tune, track, eval, Codabench upload, and collection targets. |
| benchmark/README.md | Documents required dataset layout, workflow steps, and Codabench token setup. |
| benchmark/.gitignore | Ignores local benchmark data/output artifacts. |
| benchmark/scripts/datasets.py | Centralizes dataset splits, paths, and Codabench IDs used by the workflow. |
| benchmark/scripts/data_check.py | Verifies expected dataset assets exist under DATA_ROOT. |
| benchmark/scripts/prep_data.py | Flattens vendor detections/GT into per-sequence MOT .txt files under benchmark_prep/. |
| benchmark/scripts/track_split.py | Runs a selected tracker over prepared detections and writes MOT prediction files. |
| benchmark/scripts/mot_format.py | Normalizes and packages predictions into Codabench-compatible submission zips (incl. MOT17 triplication/stubs). |
| benchmark/scripts/codabench_submit.py | Uploads/polls Codabench submissions and optionally writes a JSON summary of results. |
| benchmark/scripts/collect.py | Aggregates per-dataset JSON scores into a markdown table + summary JSON. |
| benchmark/scripts/align_mot17_val_gt.py | Filters MOT17 val GT to match the frame range covered by the YOLOX val detections. |
Comment on lines
+39
to
+41
| from trackers.core.base import BaseTracker | ||
| from trackers.tune.tuner import _run_tracker_on_detections | ||
|
|
Collaborator
Author
There was a problem hiding this comment.
mmhh, we could make them public or still use them like this
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…com/roboflow/trackers into feat/benchmark-codabench-submission
Aligns the benchmark with the train+val+test methodology: optimize on val, score on Codabench test. Co-authored-by: Cursor <cursoragent@cursor.com>
…ing. Include cbiou in COMPARISON_TRACKERS for tune/benchmark/collect workflows, and retry submission polling on transient DNS and connection errors. Co-authored-by: Cursor <cursoragent@cursor.com>
Bring in C-BIoU tracker from develop and align DanceTrack comparison notes with val-tune / Codabench test scoring. Co-authored-by: Cursor <cursoragent@cursor.com>
Add the optional trackers[reid] extra (git-pinned roboflow-reid during review) plus the lazy trackers._reid boundary that resolves reid.ReIDModel on demand. Ship association-only, torch-free modules under trackers.core.reid: the ReIDEncoder protocol, FeatureBank (per-track EMA), and appearance_similarity / extract_detection_embeddings. The model stack (encoder, weights, preprocessing, catalog, gallery eval) lives in the standalone reid package and is never vendored. Co-authored-by: Cursor <cursoragent@cursor.com>
Add appearance-IoU fusion (botsort/fusion.py) and wire it into BoT-SORT's
first and unconfirmed association stages, gated by proximity (standard IoU)
and appearance thresholds. Tracklets gain an optional per-track feature
bank; matched high-confidence detections update it. reid_model,
reid_ema_alpha, appearance_threshold, and proximity_threshold are new
BoTSORTTracker params; reid_model is excluded from CLI reflection.
Add --tracker.reid.{enable,model,device,architecture} CLI flags routed
through trackers._reid to reid.ReIDModel.from_pretrained. Importing BoT-SORT
stays torch-free (asserted by an isolation test). Add association/fusion/CLI
unit tests and a reid-backed integration smoke; docs and mkdocs nav cover
the association-only surface and link the model/eval stack out to reid.
CI installs the reid extra (unfrozen while the dep is a git pin).
Co-authored-by: Cursor <cursoragent@cursor.com>
- FeatureBank.update now L2-normalizes the incoming embedding and the resulting EMA, keeping the stored feature on the unit hypersphere (matches upstream BoT-SORT STrack.update_features). reid returns raw embeddings; normalization happens in the bank. Docstrings/tests updated. - Vectorize appearance._l2_normalize_rows. Co-authored-by: Cursor <cursoragent@cursor.com>
Load frames by MOT index, use per-sequence FPS, re-encode to H.264 for notebook/Colab playback, and embed at the combined panel width. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Inline the optional reid import in the CLI like detection/tune, and keep all BoT-SORT ReID coverage in one test module. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Default proximity to the association matrix so callers only pass a separate standard-IoU gate when association uses GIoU/DIoU/CIoU. Co-authored-by: Cursor <cursoragent@cursor.com>
Match upstream BoT-SORT: one IoU compute, proximity mask from the pre-fusion matrix, score fusion only for association. Recompute only for IoU variants. Co-authored-by: Cursor <cursoragent@cursor.com>
Only compute standard IoU for ReID proximity when association uses a variant metric; plain IoU reuses the association matrix. Co-authored-by: Cursor <cursoragent@cursor.com>
One helper builds score-fused geometry plus optional appearance fusion, shared by first and unconfirmed association. Drop thin wrappers and redundant comments. Co-authored-by: Cursor <cursoragent@cursor.com>
Drop the duplicate prerequisite helper and reject --tracker.reid.architecture without --tracker.reid.model so bare architecture cannot load random weights. Co-authored-by: Cursor <cursoragent@cursor.com>
Use the published package name `reid` (git@main until PyPI), keep unit tests free of `--extra reid`, and clean up ReID association glue/docs/tests for the cutover. Co-authored-by: Cursor <cursoragent@cursor.com>
Drop redundant embedding validation, dead reset/is_initialized API, and defensive getattr; keep device="auto" for reid to resolve. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…rkflow. Bring standalone reid package integration (BoT-SORT appearance association) onto the submission/Makefile branch so MOT17 Codabench runs can use official FastReID SBS weights.
Pass REID_ENCODER / APPEARANCE_THRESHOLD through Makefile to track_split, map MOT17 det stems to FRCNN frame folders, and install trackers[tune,reid]. Co-authored-by: Cursor <cursoragent@cursor.com>
Borda
reviewed
Aug 8, 2026
Member
There was a problem hiding this comment.
seem like oout of the scope for this PR
Borda
reviewed
Aug 8, 2026
Member
There was a problem hiding this comment.
Maybe we can simplify it now that we have benchamrks in CLI
trackers benchmark mcbyte ...
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Introduces benchmark/ directory with makefile for automatic benchamarking on MOT17, SportsMOT, and DanceTrack. Eval uses workaround for per-tracker CLI parameters (workaround to what was mentioned that would be fixed with CLI refactor). Soccernet is supported with local evaluation.
Type of Change
Testing
Checklist
Additional Context