LeRF: Learning Reference
Coordinate Frames for
Perspective Taking Reasoning

1 University of Illinois Urbana-Champaign2 Zhiyuan College, Shanghai Jiao Tong University3 University of California, Merced4 Shanghai Jiao Tong University5 Amazon AGI

✉ Corresponding author

arXiv Preprint

Perspective taking with LeRF

A girl sitting in an armchair, with a teddy bear to the left of the image. FrontLeftUp Original image
FrontLeftUp
TOOL CALL

The model first identifies the requested viewpoint.

Predicted coordinates will be passed to the renderer.

QUESTION

Where is the teddy bear from the little girl’s perspective?

Follow the steps to see the answer.
1 / 4

LeRF predicts its own frame of reference, then reasons over the rendered image. Camera-relative questions can skip frame construction.

Abstract

Perspective taking requires interpreting spatial relations from a specified viewpoint, yet vision-language models often default to the camera view. LeRF learns to construct and use explicit reference coordinate frames. Given an image and a query, the model decides whether a frame is needed, grounds the reference entity, and predicts an origin with projected front, left, and up axes. A lightweight renderer overlays the frame onto the image, allowing the same model to reason over its own visual cues. Supervised frame prediction followed by reinforcement learning improves perspective-taking performance across three benchmarks, without external perception models or explicit 3D reconstruction at inference time.

Three viewpoints—egocentric, allocentric, and hypothetical—and an example where reference-frame overlays help distinguish a worker’s left from the image’s left.
Figure 1. An entity-centered frame turns an implicit viewpoint shift into visible directional cues. The expert-overlay baseline motivates LeRF, which learns to predict its own frames. Click figure to enlarge ↗

Learning to construct and use reference frames

LeRF predicts a projected frame directly on the image, without an external perception model or a full 3D reconstruction.

LeRF inference selects frame rendering or direct reasoning. Training first supervises frame prediction, then uses GRPO with final-answer correctness rewards.
Figure 2. Selective frame construction at inference; supervised fine-tuning followed by reinforcement learning during training. Numerical frame coordinates are hidden in the reasoning turn, leaving the rendered visual cues.
STAGE I · SUPERVISED FINE-TUNING

Supervised frame prediction

Object and human pose data teach frame prediction. No-tool examples teach the model when the original image is sufficient.

STAGE II · REINFORCEMENT LEARNING

Frame-guided reasoning

GRPO rewards final-answer correctness on spatial VQA.

Perspective-taking results

LeRF-9B improves over Qwen3.5-9B on all six evaluated tasks that require a viewpoint change.

+6.91pp

Another entity’s viewpoint

OmniSpatial-PT · Allocentric

+11.08pp

An imagined viewpoint

OmniSpatial-PT · Hypothetical

+8.72pp

Person-relative direction

ViewSpatial-Bench · Relative Direction

Benchmark results

Accuracy (%) ↑
Proprietary models, selected open-source baselines, and LeRF. Bold values indicate the highest score among the displayed open-source models, including LeRF.
ModelOmniSpatial-PT3DSRBenchViewSpatial-Bench
EgoAlloHypoOrientationMulti-ObjectP - Obj. ViewP - Rel. Direction
Proprietary models
GPT-5.6-Luna (medium)83.3349.7345.7860.0455.3446.9970.07
GPT-5.6-Terra (medium)81.3755.8553.0163.3256.3945.0877.20
Claude Sonnet 5 (medium)80.3942.5549.4034.9444.7751.3151.43
Claude Sonnet 5 (high)84.3148.1445.7843.1546.8151.5160.10
Selected open-source models & LeRF
InternVL3.5-8B67.8434.2040.2427.7038.3357.6342.40
GLM-4.6V-Flash74.3134.5239.5238.3244.1360.3643.94
SpatialReasoner40.3935.1135.6652.0550.6442.3745.61
Qwen3.5-4B74.7142.5544.3442.2843.2651.0157.43
LeRF-4B72.3549.3646.7545.8844.5556.2667.85
Qwen3.5-9B80.2047.1344.5848.1748.6656.2365.51
LeRF-9B74.3154.0455.6653.7650.2961.9174.23

Selected comparisons from the paper. Bold: best score among displayed open-source methods, including LeRF; proprietary models are listed separately. All models use thinking except SpatialReasoner. Ego / Allo / Hypo denote camera, entity-centered, and imagined viewpoints. ViewSpatial-Bench columns use the person perspective.

The trade-off: camera-view accuracy decreases from 80.20% to 74.31% on OmniSpatial-PT Ego, even as performance improves on viewpoint-changing tasks.

SELECTIVE TOOL USE

Selective tool invocation

Tool invocation rates on OmniSpatial-PT: LeRF-4B uses the tool for 10.7% of Ego, 97.9% of Allo, and 97.3% of Hypo questions. LeRF-9B rates are 2.9%, 86.4%, and 84.8%, respectively.
Tool invocation rates on OmniSpatial-PT. Ego, Allo, and Hypo denote camera, entity-centered, and imagined viewpoints.

LeRF-9B invokes the renderer for just 2.94% of egocentric questions, versus over 84% of allocentric and hypothetical questions.

CONTROLLED PERTURBATION

Sensitivity to frame directions

Rotating rendered frames 90 degrees counterclockwise reduces LeRF-9B accuracy by 3.92 percentage points overall, 0.98 on Ego, 3.83 on Allo, and 7.95 on Hypo.
LeRF-9B on OmniSpatial-PT with original frames and frames rotated 90° counterclockwise at evaluation time.

Rotating rendered frames by 90° reduces hypothetical accuracy by 7.95 percentage points, indicating that the model uses the directional cues.

Qualitative examples

LeRF answers directly for a camera-relative query, and constructs reference frames for entity-centered and imagined viewpoints.

Three qualitative examples: a camera-view answer of below without tool use; a teddy bear to a girl's right using her frame; and a door to an imagined seated observer's right.
Figure 3. Examples of selective tool use and frame-guided reasoning. Reference-object boxes are illustrative and are not included in the model inputs.

Additional comparisons with baselines

Appendix C.3 compares LeRF with SpatialReasoner, APC + Qwen3.5-9B, and Qwen3.5-9B. These examples show how a camera-relative answer can differ from the requested entity-centered answer.

Appendix C.3: three baseline comparisons. LeRF answers right from the person with red gloves, counts two people to the cat’s right, and places the window on the projection wall’s right. Baselines give camera-relative answers or cannot determine the answer.
Figure 4. Additional qualitative comparisons from Appendix C.3. LeRF constructs an entity-centered reference frame in each example. Click to enlarge.

(a) A person’s viewpoint. All three baselines answer “left” in the camera view. LeRF grounds the person with red gloves and answers “right” from that person’s perspective.

(b) A cat’s viewpoint. The baselines count one person. LeRF makes the cat-centered right direction explicit and correctly counts two.

(c) A wall’s orientation. The baselines answer “left” or cannot determine the answer. LeRF uses the projection wall’s intrinsic orientation and identifies the window on its right.

Scope & limitations

LeRF currently focuses on static images. Occlusion, visual ambiguity, and uncertain entity orientation can affect predicted frames and propagate to answers. Tracking changing reference frames in dynamic scenes remains future work.

Paper figure