Gorio Tech Blog search

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs 요약 설명

|

목차

이번 글에서는 RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs 논문의 핵심 포인트만 간단히 정리한다.

영문판 보기


요약

  • RelateAnything는 객체 클래스 라벨을 모델 입력으로 사용하지 않고, 이미지의 순서가 있는 영역 쌍과 추론 시점에 주어진 predicate 문자열 목록을 점수화하는 relation prediction 문제로 정의한다. 논문은 53M-parameter 모델, RA-4M corpus, OV-SGG-Bench를 제안한다. 컴파일된 배포 경로는 A40에서 프레임당 20 ms를 기록했고, eager end-to-end 비교는 25.0 ms를 보고한다.

개요는 3개 산출물을 연결한다. RA-4M은 검증된 free-text relation을 제공하고, RelateAnything는 pixel, region, predicate string을 분리하며, OV-SGG-Bench는 6개 axis를 평가한다. RelateAnything의 보고된 OVS composite 40.1과 OvSGTR의 11.8도 보여준다.

corpus, relation model, 6축 평가 protocol의 pipeline 개요
corpus, relation model, 6축 평가 protocol의 pipeline 개요

3 The RelateAnything model

이미지 I와 영역 B = {b_i}_{i=1}^N가 주어지면, 모델은 p가 추론 시 제공되는 vocabulary V에 속하는 점수화된 triplet (i, j, p)를 반환한다. 객체 category는 명시된 training loss에서만 사용하며 forward pass에서는 사용하지 않는다. predicate는 고정된 학습 classifier 대신 text embedding으로 표현한다.

3.2 Architecture

fine-tuned DINOv3 backbone은 융합된 dense feature를 만들고, coordinate-aware query는 순서가 있는 쌍의 subject, object, union, contact-region feature를 pooling한다. Relation Transformer와 deformable scene read가 이 쌍 표현을 정제한다. spatial branch와 semantic branch가 frozen predicate embedding에 대해 점수를 계산하고, predicate-conditioned gate가 2개의 cosine channel을 혼합한다.

3.3 Training with 10⁴ predicates

학습에는 19,103-string predicate bank와 synonym group을 positive로 사용하는 batch-local InfoNCE를 사용한다. 주석되지 않은 predicate도 실제로 성립할 수 있으므로, positive-unlabeled supervision에서는 함께 주석될 가능성이 높은 negative의 가중치를 낮추고 directional inverse는 완전한 가중치의 negative로 유지한다. distilled 512-dimensional text encoder는 synonym cosine 0.71을 유지하면서 mean inverse-pair cosine을 0.92에서 0.09로 낮춘다. positive에 대해서는 이 글을 참조하라.

teacher text space에서는 inverse spatial predicate가 synonym만큼 가깝게 배치되지만, distilled student는 synonym group을 유지하면서 inverse direction을 분리한다. 이는 visual relation alignment 전에 text-space distillation을 수행하는 근거를 제시한다.

spatial predicate의 teacher 및 distilled text-embedding geometry
spatial predicate의 teacher 및 distilled text-embedding geometry

4 RA-4M: a verified free-text relation corpus

RA-4M은 474,413개 이미지, 4,282,531개 relation, 10,102개의 서로 다른 free-text predicate로 구성된다. 같은 이미지와 box의 원본 annotation과 비교하면 1.7× 더 조밀하고 predicate vocabulary는 107× 더 크다. annotation을 고정된 class 목록으로 투영하지 않고 surface form을 유지한다.

4.1 Generation and geometric verification

open-weight 26B VLM은 번호가 매겨진 box marker를 대상으로 semantic, graph-growth, spatial relation을 생성한다. deterministic geometric check는 raw candidate의 11.3%를 제거하고 inverse form이 있을 때 모순된 directional role을 수정한다. gaze와 잘못된 instance grounding처럼 기하적으로 제약할 수 없는 사례는 명시적으로 남는 위험으로 인정한다.

5 OV-SGG-Bench: analysis of scene-graph evaluation and a cross-dataset protocol

OV-SGG-Bench는 기존 scene-graph recall이 visual relation prediction, corpus overlap, annotation propensity, detector setting을 함께 측정한다고 지적한다. 6개 axis는 transfer (A1), adjudicated negative에 대한 precision (A2), full-vocabulary prediction (A3), detector-mode deployment (A4), judged graph quality (A5), adversarial spatial relation (A6)을 측정한다.

5.2 Three priors rewarded by recall

oracle object category를 받은 pixel-free frequency table은 edge별 micro accuracy에서는 모델을 앞서지만, VG150, PSG, IndoorVG 전체의 macro accuracy에서는 뒤처진다. 저자들은 shared triplet mass, object-category prior, annotation propensity가 전이 가능한 visual relation prediction을 입증하지 않고도 recall을 높일 수 있는 shortcut이라고 지적한다.

5.3 Two scoring conventions with effects larger than those measured

2개의 scoring choice는 보고된 method gap보다 큰 차이를 만들 수 있다. unconstrained matching은 하나의 ground-truth object에 여러 detection이 점수를 얻도록 할 수 있고, detector threshold와 box cap은 달성 가능한 pair-recall ceiling을 바꾼다. 따라서 protocol은 graph-constrained scoring을 사용하고 detector operating point 및 그 ceiling을 보고한다.

5.4 The six axes

OVS composite는 A1, A2, A4, A5, A6에 chance correction을 적용한 뒤 harmonic mean으로 결합해 불균형한 system에 불이익을 준다. A3는 synonym-tolerant free-text matching이 scorer에 민감하므로 별도로 보고한다. 논문은 composite 하나보다 평가 축 벡터를 실질적인 평가로 다룬다.

detector weight를 고정해도 confidence threshold와 box cap이 pair-recall ceiling을 크게 바꾼다. 이 그림은 detection-mode scene-graph 결과에 명시적인 operating point와 ceiling이 필요하다는 논문의 주장을 뒷받침한다.

detector operating-point sweep에 따른 pair-recall ceiling
detector operating-point sweep에 따른 pair-recall ceiling

6 Results

RelateAnything는 OVS composite 40.1을 기록했으며, 모든 composite axis에서 실행한 유일한 baseline인 OvSGTR는 11.8을 기록했다. 보고된 향상은 transfer, explicit-negative precision, shared-detector deployment, information-weighted graph quality, SpatialSense AUC에 걸친다. 다만 논문은 judge-based 측정과 spatial 측정의 해석에 한계가 있다고 명시한다.

이 표는 핵심 6축 비교를 제시한다. RelateAnything는 나열된 각 cell에서 실행 가능한 최선의 baseline을 앞서며, 모든 composite axis에서 실행된 유일한 baseline인 OvSGTR의 11.8에 비해 OVS composite 40.1을 얻는다.

transfer, precision, open vocabulary, deployment, graph quality, spatial relation 전반의 OV-SGG-Bench 결과
transfer, precision, open vocabulary, deployment, graph quality, spatial relation 전반의 OV-SGG-Bench 결과

6.1 A1: cross-dataset transfer on four benchmarks

ground-truth region을 사용한 cross-dataset transfer에서 RelateAnything는 VG150에서 F1@50/mR@50 36.9/28.2, PSG에서 34.7/30.6, IndoorVG에서 37.8/29.5, zero-shot HICO-DET tower에서 18.7/12.7을 달성한다. OvSGTR 대비 mean-recall margin은 2.3–3.5×이고 rare-predicate recall margin은 5–21×이다. baseline은 ground-truth object label을 입력으로 받는다.

transfer 표는 ground-truth region 조건에서 VG150, PSG, IndoorVG, HICO-DET 결과를 보고한다. zero-shot HICO-DET tower와 5% HICO-DET relation share로 학습한 버전을 구분하고, baseline이 ground-truth object label을 사용함을 기록한다.

VG150, PSG, IndoorVG, HICO-DET의 cross-dataset transfer
VG150, PSG, IndoorVG, HICO-DET의 cross-dataset transfer

6.2 A2: precision on adjudicated negatives

Haystack의 2,870개 positive와 23,174개 explicit negative에서 모델은 mean federated AP를 52.1에서 72.6으로, rare-predicate fAP를 44.6에서 70.7로, P-AUC를 83.9에서 93.7로 높인다. sampler는 baseline의 100%와 달리 labelled cell의 89.0%를 포괄한다. 포함되지 않은 cell은 제외하지 않고 0점을 부여한다.

6.4 A4: detection-mode graphs on a shared detector

PSG에서 confidence 0.05와 이미지당 최대 60개 box를 사용하는 shared YOLO-World detector 조건에서 RelateAnything는 R@50 25.4, mR@50 20.5, rare recall 17.0, wR@50 20.0을 기록한다. OvSGTR는 각각 17.7, 6.2, 1.2, 4.0을 기록한다. 두 system은 같은 69.6 pair-recall ceiling의 제약을 받으므로 이 detector configuration에서 relation prediction을 분리해 비교한다.

이 표는 Haystack precision, full-bank open-vocabulary ranking, shared-detector PSG 결과를 상세히 제시한다. pair coverage와 공유된 69.6 pair-recall ceiling을 명시해 비교에서 통제한 요소와 통제하지 않은 요소를 분명히 한다.

precision, full-vocabulary prediction, detector-mode scene graph의 상세 결과
precision, full-vocabulary prediction, detector-mode scene graph의 상세 결과

6.5 A5: graph quality judged without ground truth

A5에서 Qwen3-VL-8B는 ground truth 없이 개별 assertion을 판단하고, 수용된 relation은 predicate surprisal로 가중한다. 같은 graph depth에서 RelateAnything는 이미지당 18.6 true bit를 얻어 OvSGTR의 13.4보다 높다. 그러나 judge는 이미지당 baseline assertion을 약간 더 많이 수용한다. 논문은 whole-graph 판정과 per-relation 판정이 baseline에 유리할 수 있다는 점도 별도로 보고한다.

같은 output depth에서 baseline은 judge가 수용한 relation이 약간 더 많지만, RelateAnything가 수용된 predicate의 평균 surprisal이 더 높아 이미지당 true bit가 더 많다. 예시는 baseline이 predicate on에 집중하는 현상도 보여준다.

information-weighted graph-quality 분해와 graph 예시
information-weighted graph-quality 분해와 graph 예시

6.6 A6: spatial relations against adversarial negatives

SpatialSense에서 RelateAnything는 macro AUC 0.690과 pooled AUC 0.675를 기록했으며, 같은 cell에서 OvSGTR는 0.591과 0.617을 기록했다. 저자들은 이 결론을 명시적으로 제한한다. benchmark의 boxes-only baseline이 68.8%에 이르므로 AUC 약 0.69는 일부 prior에 대한 저항을 뜻할 뿐, 결정적인 visual spatial reasoning을 입증하지는 않는다.

7 Analysis: decoupling object and relation prediction

분석은 OvSGTR이 object label에 더 크게 의존함을 보인다. object-label-pair lookup은 그 출력의 88–90%를 재현하지만, label을 전혀 받지 않는 RelateAnything에서는 69–70%를 재현한다. lesion 분석은 제안 모델의 semantic-logit variance 중 87–93%를 pair context에, 약 0.1%와 0.0%를 subject와 object identity feature에 귀속한다. representation probe는 이 architecture와 training recipe에 한정된다.

label-dependence 분석은 OvSGTR prediction이 object-label-pair lookup table로 더 잘 재현되고 label을 섞어도 대부분 변하지 않음을 보여준다. RelateAnything는 label을 관찰할 수 없고 residual predicate entropy가 더 높으며, 이 평가에서 50개 predicate를 모두 사용한다.

object-label pair에 대한 output dependence
object-label pair에 대한 output dependence

8 Analysis: in-domain versus cross-dataset measurement

corpus를 2배로 늘리면 in-domain R@50과 mR@50은 15%와 56% 향상되지만, zero-shot transfer는 각각 약 2.5–3.8%와 9.9–11.9%만 향상된다. 저자들은 in-domain scaling slope가 transfer gain을 대략 5× 과대평가한다고 추정한다. sparse Open Images relation supervision을 추가하면 자신들의 training setup에서 측정한 모든 development axis가 저하됐다고 보고한다.

그림은 matched-step in-domain scaling과 zero-shot transfer를 비교한다. corpus를 2배로 늘릴 때 in-domain에서 측정한 이득이 transfer 이득을 상당히 과대평가한다는 주장의 시각적 근거를 제시한다.

in-domain 및 zero-shot transfer에서의 corpus scaling gain 비교
in-domain 및 zero-shot transfer에서의 corpus scaling gain 비교

10 Discussion and limitations

eager A40 비교에서 RelateAnything와 YOLO-World는 25.0 ms 및 40.0 FPS를, OvSGTR Swin-T는 194.0 ms 및 5.1 FPS를 기록한다. 이는 서로 다른 box budget, resolution, precision, pair count를 사용하는 deployed-configuration 비교다. 한계로는 human-audited RA-4M precision 추정 부재, 모든 benchmark에 걸친 일반적인 held-out-concept 연구 부재, 여러 single-seed 결과, detector-dependent pair recall, A5에서 baseline보다 낮은 per-relation acceptance가 있다.

latency 표는 architecture-controlled system이 아니라 deployed configuration을 비교한다. 더 적은 box와 scored pair, 더 낮은 resolution, bf16을 사용하는 RelateAnything가 보고된 protocol에서 A40 median latency 25.0 ms를 기록하며, OvSGTR Swin-T는 194.0 ms를 기록한다.

OvSGTR과 비교한 end-to-end inference cost
OvSGTR과 비교한 end-to-end inference cost

부록

  • appendix는 architecture, objective, corpus generation과 verification, calibration, benchmark reproduction, 추가 분석, backbone scaling, ablation, inference measurement을 명시한다. 보조 세부 사항으로는 1.3% seed-matched noise floor, 혼합 annotation scope를 위한 source-aware negative masking, composite 성능이 더 큰 backbone에서 단조롭게 개선되지 않아 선택한 released ViT-S/16+ tower가 있다.

짧은 생각

논문은 open-vocabulary relation interface를 free-text machine-generated supervision 및 multi-axis evaluation protocol과 결합한다. 가장 강한 방법론적 기여는 recall 결과에 transfer setting, explicit-negative precision, matching convention, detector operating point, pair-recall ceiling을 함께 제시해야 한다는 주장이다. corpus precision과 VLM-judge 의존성은 여전히 해결되지 않은 우려 사항이다.