Language-conditioned object highlighting for simulated prosthetic vision
Abstract
Retinal prostheses offer a promising avenue for partial vision restoration, yet their utility remains severely hindered by the perceptual limitations of current implants. Existing pipelines often attempt a global reconstruction of the visual scene; however, given the extremely low information bandwidth of these devices, this approach inevitably sacrifices critical detail, causing small or distant objects to vanish. While recent paradigms have explored scene simplification, they often rely on rigid saliency maps, lack realistic egocentric dynamics, or neglect non-linear distortions like axonal deformations. To address these challenges, we propose a modular framework that bridges open-vocabulary semantic understanding with biophysical phosphene optimization. Our pipeline leverages vision-language and segmentation foundation models to dynamically isolate target objects from natural language prompts. This object-conditioned visual target is then processed by a bio-inspired phosphene autoencoder, trained with a novel region-weighting strategy that optimizes electrical stimuli to maximize the perceptual fidelity of user-prioritized targets. In simulation, our framework significantly improves the preservation of small, task-relevant objects across varying shapes and scales, outperforming global reconstruction and heuristic baselines in both perceptual fidelity and geometric alignment.