Special Issue on Frontier Technologies and Applications of Computer Vision
Published 21 August, 2026
Introduction:
Computer vision is evolving from "perception and recognition" through "reasoning and generation" toward "physical interaction." Foundation models, represented by vision-language models (VLMs) and visual object models (VOMs), endow machines with strong cross-modal semantic understanding and scene-decomposition capabilities. Artificial intelligent-generated content AIGC-driven visual generation is reconstructing the physical world with genuine physical logic and spatial consistency, with 3D/4D geometric vision leaping from static reconstruction to dynamic spatiotemporal modeling. Meanwhile, vision-language-action (VLA) models close the loop from perception to decision-making and execution, opening new pathways for embodied intelligence and autonomous driving. This special issue aims to showcase the latest theoretical breakthroughs, key technological advances, and representative application outcomes in this rapidly evolving field.
Topics covered:
(1) (VLMs)
- Cross-modal semantic alignment, unified representation learning, and multimodal fusion
- Efficient pre-training, parameter-efficient fine-tuning, and prompt engineering
- Enhanced visual reasoning, hallucination mitigation, and interpretability
- Lightweight architecture design and on-device deployment
(2) Visual Object Models (VOMs)
- Object-centric scene decomposition, unsupervised object discovery, and representation learning
- Visual value models and inference-time search optimization
- Object permanence and occlusion reasoning in dynamic scenes
- Open-world object detection and recognition
(3) 3D/4D Geometric Vision
- Efficient reconstruction with Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS)
- Dynamic scene-flow estimation and 4D spatiotemporal modeling
- Multi-view geometry and structured-light / Time-of-Flight (ToF) depth sensing
- Differentiable rendering and inverse graphics
(4) Visual Artificial Intelligent-Generated Content (AIGC) and Generative Models
- Diffusion models and their applications in image / video generation
- Controllable generation, conditional editing, and style transfer
- Learning of physical laws and world simulation in video generation
- AIGC-based data synthesis to assist visual model training
(5) Vision–Language–Action Models (VLAs)
- End-to-end embodied decision-making and action tokenization
- Sim-to-Real transfer and domain adaptation
- Applications of VLAs in robotic manipulation and autonomous driving
- Closed-loop perception–decision–control systems
(6) Cross-disciplinary Integration and System Implementation
- Scene understanding and reasoning through coordinated VLM + VOM
- Data generation and augmentation methods for visual model training
- End-to-end joint optimization of optical imaging and visual models
- Applications in vertical scenarios such as embodied intelligence, autonomous driving, medical imaging, and remote sensing
Important deadlines:
- Date First Submission Expected: 1 January 2027
- Final Manuscript Submission Deadline: 31 July 2027
- Editorial Acceptance Deadline: 31 October 2027
Guest editors:
- Xuelong Li, China Telecom AI Research Institute (TeleAI), Research Interests: Optoelectronic Imaging and Intelligent Information Processing
- Zhe Sun, Northwestern Polytechnical University, Research Interests: Underwater Optical Detection and Imaging; sunzhe@nwpu.edu.cn
- Hongyuan Zhang, The University of Hong Kong, Research Interests: Positive Excitation Noise: Characterization, Understanding, and Generation; hyzhang98@gmail.com
- Tao Chen, Fudan University, Research Interests: Computer vision, (Multidimensional) Data Analytics, Machine Learning and Pattern Recognition; eetchen@fudan.edu.cn
- Jingchun Zhou, Dalian Maritime University, Research Interests: Computer Vision, Image Enhancement and Restoration, Image Quality Assessment; zhoujingchun@dlmu.edu.cn
- Tong Tian, Friedrich Schiller University Jena, Research Interests: Development of versatile single-pixel imaging technologies based on the synergistic integration of physical models and computational algorithms. Research encompasses optical system design, physics-informed optimization, and deep learning to achieve robust and efficient imaging in computational ghost imaging, phase imaging, and other challenging optical scenarios. tong.tian@uni-jena.de