RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs | Summary
11 Sep 2026 | Paper Review Open-Vocabulary Scene Graph Generation Relation Prediction Positive-Unlabeled LearningContents
- Summary
- 3 The RelateAnything model
- 4 RA-4M: a verified free-text relation corpus
- 5 OV-SGG-Bench: analysis of scene-graph evaluation and a cross-dataset protocol
- 6 Results
- 7 Analysis: decoupling object and relation prediction
- 8 Analysis: in-domain versus cross-dataset measurement
- 10 Discussion and limitations
- Appendix
- Brief Thoughts
This article explains the key points of RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs.
Summary
- RelateAnything formulates relation prediction as scoring ordered region pairs from an image against an inference-time list of predicate strings, without using object class labels as model inputs. The paper introduces a 53M-parameter model, the RA-4M corpus, and OV-SGG-Bench; its compiled deployment path is reported at 20 ms per frame on an A40, while the eager end-to-end comparison reports 25.0 ms.
The overview links the three artifacts: RA-4M supplies verified free-text relations, RelateAnything separates pixels, regions, and predicate strings, and OV-SGG-Bench evaluates six axes. It also shows the reported OVS composite of 40.1 for RelateAnything and 11.8 for OvSGTR.
3 The RelateAnything model
Given an image I and regions B = {b_i}_{i=1}^N, the model returns scored triplets (i, j, p), where p belongs to an inference-supplied vocabulary V. Object categories are used only in specified training losses, not in the forward pass, and predicates are represented by text embeddings rather than a fixed learned classifier.
3.2 Architecture
A fine-tuned DINOv3 backbone yields fused dense features, and coordinate-aware queries pool subject, object, union, and contact-region features for ordered pairs. A Relation Transformer and deformable scene read refine these pair representations; spatial and semantic branches score them against frozen predicate embeddings, and a predicate-conditioned gate mixes the two cosine channels.
3.3 Training with 10⁴ predicates
Training uses a 19,103-string predicate bank and batch-local InfoNCE with synonym groups as positives. Because unannotated predicates may still hold, likely co-annotated negatives are discounted under positive-unlabeled supervision while directional inverses remain full-weight negatives; a distilled 512-dimensional text encoder reduces mean inverse-pair cosine from 0.92 to 0.09 while retaining synonym cosine 0.71.
The teacher text space places inverse spatial predicates nearly as close as synonyms, whereas the distilled student preserves synonym groupings while separating inverse directions. This supports the motivation for text-space distillation before visual relation alignment.
4 RA-4M: a verified free-text relation corpus
RA-4M contains 474,413 images, 4,282,531 relations, and 10,102 distinct free-text predicates. Relative to the source annotations on the same images and boxes, it is 1.7× denser, has 107× the predicate vocabulary, and retains surface forms rather than projecting annotations into a fixed class list.
4.1 Generation and geometric verification
An open-weight 26B VLM generates semantic, graph-growth, and spatial relations over numbered box markers. Deterministic geometric checks reject 11.3% of raw candidates, repair contradicted directional roles when an inverse form exists, and leave geometrically unconstrained cases—such as gaze and wrong-instance grounding—as explicitly acknowledged residual risks. For VLM, see this post.
5 OV-SGG-Bench: analysis of scene-graph evaluation and a cross-dataset protocol
OV-SGG-Bench argues that conventional scene-graph recall mixes visual relation prediction with corpus overlap, annotation propensity, and detector settings. Its six axes measure transfer (A1), precision on adjudicated negatives (A2), full-vocabulary prediction (A3), detector-mode deployment (A4), judged graph quality (A5), and adversarial spatial relations (A6).
5.2 Three priors rewarded by recall
A pixel-free frequency table supplied with oracle object categories exceeds the model on per-edge micro accuracy, but trails it on macro accuracy across VG150, PSG, and IndoorVG. The authors identify shared triplet mass, object-category priors, and annotation propensity as shortcuts that can improve recall without establishing transferable visual relation prediction.
5.3 Two scoring conventions with effects larger than those measured
Two scoring choices can exceed reported method gaps: unconstrained matching can credit multiple detections for one ground-truth object, and detector thresholds and box caps alter the attainable pair-recall ceiling. The protocol therefore uses graph-constrained scoring and reports detector operating points and their ceilings.
5.4 The six axes
The OVS composite chance-corrects A1, A2, A4, A5, and A6, then combines them by a harmonic mean to penalize uneven systems. A3 is reported separately because synonym-tolerant free-text matching is sensitive to the scorer; the paper treats the six-axis vector, rather than the composite alone, as the substantive evaluation.
With detector weights fixed, confidence thresholds and box caps substantially alter pair-recall ceilings. The figure supports the paper's argument that detection-mode scene-graph results require explicit operating points and ceilings.
6 Results
RelateAnything obtains an OVS composite of 40.1, compared with 11.8 for OvSGTR, the only baseline run on all composite axes. The reported gains span transfer, explicit-negative precision, shared-detector deployment, information-weighted graph quality, and SpatialSense AUC, although the paper qualifies the interpretation of judge-based and spatial measurements.
This is the main six-axis comparison. RelateAnything exceeds the best runnable baseline in each listed cell and obtains a 40.1 OVS composite against 11.8 for the single baseline, OvSGTR, that is run across all composite axes.
6.1 A1: cross-dataset transfer on four benchmarks
On cross-dataset transfer with ground-truth regions, RelateAnything reaches F1@50/mR@50 of 36.9/28.2 on VG150, 34.7/30.6 on PSG, 37.8/29.5 on IndoorVG, and 18.7/12.7 for the zero-shot HICO-DET tower. Against OvSGTR, its mean-recall margins are 2.3–3.5× and its rare-predicate recall margins are 5–21×, while the baseline receives ground-truth object labels.
The transfer table reports results on VG150, PSG, IndoorVG, and HICO-DET with ground-truth regions. It separates the zero-shot HICO-DET tower from the version trained with a 5% HICO-DET relation share and documents the baseline's use of ground-truth object labels.
6.2 A2: precision on adjudicated negatives
On Haystack’s 2,870 positives and 23,174 explicit negatives, the model improves mean federated AP from 52.1 to 72.6, rare-predicate fAP from 44.6 to 70.7, and P-AUC from 83.9 to 93.7. Its sampler covers 89.0% of labelled cells rather than the baseline’s 100%, and uncovered cells are assigned a zero score rather than excluded.
6.4 A4: detection-mode graphs on a shared detector
With a shared YOLO-World detector on PSG at confidence 0.05 and at most 60 boxes per image, RelateAnything records R@50 25.4, mR@50 20.5, rare recall 17.0, and wR@50 20.0; OvSGTR records 17.7, 6.2, 1.2, and 4.0. Both systems are bounded by the same 69.6 pair-recall ceiling, isolating relation prediction under this detector configuration.
The table details Haystack precision, full-bank open-vocabulary ranking, and shared-detector PSG results. It makes pair coverage and the shared 69.6 pair-recall ceiling explicit, clarifying what is and is not controlled in the comparisons.
6.5 A5: graph quality judged without ground truth
For A5, Qwen3-VL-8B judges individual assertions without ground truth, and accepted relations are weighted by predicate surprisal. At matched graph depth, RelateAnything obtains 18.6 true bits per image versus 13.4 for OvSGTR, although the judge accepts slightly more baseline assertions per image; the paper separately reports that whole-graph and per-relation verdicts can favor the baseline.
At matched output depth, the baseline has slightly more judge-accepted relations, but RelateAnything's accepted predicates have greater average surprisal and hence more true bits per image. The examples also expose the baseline's concentration on the predicate on.
6.6 A6: spatial relations against adversarial negatives
On SpatialSense, RelateAnything reaches macro AUC 0.690 and pooled AUC 0.675, compared with OvSGTR’s 0.591 and 0.617 on the same cells. The authors explicitly limit this conclusion: the benchmark’s boxes-only baseline reaches 68.8%, so an AUC near 0.69 indicates resistance to some priors rather than conclusive visual spatial reasoning.
7 Analysis: decoupling object and relation prediction
The analysis finds greater object-label dependence in OvSGTR: object-label-pair lookup reproduces 88–90% of its outputs, versus 69–70% for RelateAnything, which never receives labels. Lesions attribute 87–93% of the proposed model’s semantic-logit variance to pair context and about 0.1% and 0.0% to subject and object identity features; its representation probes are specific to this architecture and training recipe.
The label-dependence analysis shows that OvSGTR predictions are more reproducible from object-label-pair lookup tables and remain largely unchanged when labels are shuffled. RelateAnything cannot observe labels, has higher residual predicate entropy, and uses all 50 predicates in this evaluation.
8 Analysis: in-domain versus cross-dataset measurement
Doubling the corpus improves in-domain R@50 and mR@50 by 15% and 56%, but improves zero-shot transfer by only about 2.5–3.8% and 9.9–11.9%, respectively. The authors estimate that in-domain scaling slopes overstate transfer gains by roughly 5×, and report that adding sparse Open Images relation supervision degraded every measured development axis under their training setup.
The figure compares matched-step in-domain scaling with zero-shot transfer. It is the visual evidence for the claim that gains measured in-domain substantially overstate the transfer benefit of doubling the corpus.
10 Discussion and limitations
The eager A40 comparison reports 25.0 ms and 40.0 FPS for RelateAnything with YOLO-World, versus 194.0 ms and 5.1 FPS for OvSGTR Swin-T; this is a deployed-configuration comparison with different box budgets, resolution, precision, and pair counts. Limitations include no human-audited RA-4M precision estimate, no general held-out-concept study across all benchmarks, several single-seed findings, detector-dependent pair recall, and lower per-relation acceptance than the baseline in A5.
The latency table compares deployed configurations rather than architecture-controlled systems: RelateAnything uses fewer boxes, fewer scored pairs, lower resolution, and bf16. Under the reported protocol it documents a 25.0 ms A40 median latency versus 194.0 ms for OvSGTR Swin-T.
Appendix
- The appendix specifies the architecture, objective, corpus generation and verification, calibration, benchmark reproduction, additional analyses, backbone scaling, ablations, and inference measurements. Supporting details include a 1.3% seed-matched noise floor, source-aware negative masking for mixed annotation scopes, and the released ViT-S/16+ tower selected over larger backbones whose composite performance did not improve monotonically.
Brief Thoughts
The paper couples an open-vocabulary relation interface with free-text machine-generated supervision and a multi-axis evaluation protocol. Its strongest methodological contribution is the insistence that recall results be accompanied by transfer settings, explicit-negative precision, matching conventions, detector operating points, and pair-recall ceilings; corpus precision and VLM-judge dependence remain unresolved concerns.