My Account Log in

1 option

Diffusion Framework with Cross-Modality Fusion Perception for Autonomous Driving in Urban Traffic School of Management, Hefei University of Technology

SAE Technical Papers (1906-current) Available online

View online
Format:
Book
Conference/Event
Author/Creator:
Qu, Yanwei, author.
Mo, Hangjie, author.
Conference Name:
Interntional Conference on the New Energy and Intelligent Vehicles (2025-11-02 : Hefei, China)
Language:
English
Subjects (All):
Autonomous vehicles.
Electric vehicles.
Lidar.
Sensors and actuators.
Cameras.
Logistics.
Simulation and modeling.
Vehicle acceleration.
Research and development.
Local Subjects:
Autonomous vehicles.
Electric vehicles.
Lidar.
Sensors and actuators.
Cameras.
Logistics.
Simulation and modeling.
Vehicle acceleration.
Research and development.
Physical Description:
1 online resource
Place of Publication:
Warrendale, PA SAE International 2026
Summary:
End-to-end autonomous driving in urban environments faces three core challenges. First, camera and LiDAR sensor heterogeneity causes cross-modal perception inconsistencies and sensor fusion instability. Second, diffusion models suffer from training instability due to scale variance and distribution changes, which limits generalization. Third, traditional trajectory decoders lack structured interaction with semantic elements, thereby undermining planning rationality. To address these issues, CMFPNet introduces an integrated framework with three key modules. The HGCF-Backbone integrates LiDAR and camera features using channel focus, deformable cross-focus, and state space modeling to enhance semantic alignment. The NST module maps physical trajectories to normalized space, employing truncated diffusion sampling for stable generation in just 24 steps. The NDA models trajectory generation as a semantic narrative, utilizing a six-stage semantic attention flow incorporating BEV context, interactive dynamics, and self-states. Experiments on the NAVSIM dataset demonstrate CMFP Net's superiority over existing baselines, showing outstanding generalization and trajectory stability in challenging scenarios. Notably, the truncated sampling strategy achieves an 810 acceleration during inference while maintaining decision accuracy and reducing computational costs. CMFPNet provides a scalable, semantically consistent solution for diffusion-based autonomous driving with significant potential in both research and practical deployment
Notes:
Vendor supplied data
Access Restriction:
Restricted for use by site license

The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.

Find

Home Release notes

My Account

Shelf Request an item Bookmarks Fines and fees Settings

Guides

Using the Find catalog Using Articles+ Using your account