Abstract
Scene coordinate regression (SCR) is a widely studied approach in visual localization, gaining increasing attention for its ability to predict 2D-3D correspondences. Most SCR methods assume that all regions within an image contribute equally to learning these correspondences, or prioritize semantically salient areas such as buildings. Yet in practice, only a subset of regions offers meaningful cues for accurate 2D-3D correspondence. Our study reveals a clear mismatch between these human-defined semantic expectations and the regions that are empirically most useful for SCR. By projecting known generated 3D points map onto 2D images and comparing them with text embeddings, we observe that certain concept words, which are not typically prioritized by human intuition, exhibit stronger alignment with triangulated 3D points than broader semantic categories like buildings. While text-guided hard masking based on these more aligned prompts can enhance performance in some datasets, it often leads to degraded results in others, highlighting the sensitivity of such approaches to scene-specific variability. Taken together, these findings position our work as an empirical analysis of prompt-guided masking, clarifying both its current limitations and its potential for future adaptive, data-driven approaches to region selection in SCR.
| Original language | English |
|---|---|
| Pages (from-to) | 3689-3701 |
| Number of pages | 13 |
| Journal | International Journal of Control, Automation and Systems |
| Volume | 23 |
| Issue number | 12 |
| DOIs | |
| State | Published - 2025.12 |
Keywords
- Scene coordinate regression
- semantic segmentation
- text-guided model
- visual localization
Fingerprint
Dive into the research topics of 'Are We Looking at the Right Place? Exploring Regional Bias in Prompt-guided Scene Coordinate Regression'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver