PDF(10357 KB)
PDF(10357 KB)
PDF(10357 KB)
自动驾驶地图端到端生成技术综述
End-to-end autonomous driving map generation: a survey
自动驾驶地图是支撑智能汽车安全运行的关键基础设施,尤其在道路形态复杂、路况变化频繁的国内城市道路场景,其对系统安全性和稳定性的作用更为重要。长期以来,地图生产主要依赖专业采集和人工标注,普遍受制于高成本、低效率和长更新周期,难以满足自动驾驶地图大规模覆盖和持续更新的需求。在这一背景下,端到端生成逐渐成为该领域的重要演进方向。围绕这一方向,该文从传统多阶段制图的局限出发,梳理自动驾驶地图生成技术的主要发展脉络,比较不同端到端方法在建模思路和实现路径上的差异,并进一步讨论大模型及视觉语言模型带来的新可能及其现实约束。在此基础上,该文对当前尚未解决的关键问题和后续技术演进趋势进行归纳和讨论。
Significance: Autonomous driving maps are essential for the safe operation of intelligent vehicles, especially on Chinese urban roads, where road structures are complex and traffic conditions change frequently. For a long time, map production has relied on specialized surveying and extensive manual annotation, resulting in high costs, low efficiency, and long update cycles, making large-scale coverage and continuous updating difficult to sustain. Under these conditions, end-to-end autonomous driving map generation has gradually become an important direction in this field. Its fundamental goal is to infer structured map outputs more directly from sensor observations, thereby reducing dependence on complex intermediate procedures and intensive manual intervention. With the continued development of deep learning, the two core tasks of autonomous driving map generation—lane network generation and lane topology prediction—have both shifted toward end-to-end approaches. Concurrently, industrial end-to-end systems have been developed for city-scale map production and updating, while large models, vision-language models, and agent-based systems have further expanded the research space in this area. Nevertheless, several challenges remain unresolved, including generalization in complex scenes, dynamic change detection, interpretability, and consistency in multi-vehicle collaboration. Therefore, it is important to systematically review existing research to provide clearer guidance for the future development of this field. Progress: For lane network generation, the technical route has evolved from early convolutional neural network (CNN)-based methods to Transformer-based and hybrid architectures. CNN-based methods were widely adopted at an early stage due to their mature operators, stable training, and low deployment cost. They are effective in modeling local geometry; however, their ability to preserve long-range structural consistency becomes more limited in complex scenes. Transformer-based methods later emerged as a major direction because query-based decoding is well suited for structured instance prediction and global relation modeling. Subsequent studies further expanded this line of work to include geometric constraints, map element representation, prior-guided prediction, and temporal consistency in online mapping. Hybrid architectures combine convolutional feature extraction, Transformer-based reasoning, graph modules, and temporal memory, enabling local geometric precision, topological consistency, and engineering feasibility to be addressed within a unified framework. A similar shift can be observed in lane topology prediction. Early methods mainly relied on local connection inference, whereas later studies increasingly treated topology as a structured prediction problem involving order, connectivity, and global consistency. Transformer-based methods strengthened long-range dependency modeling and gradually incorporated geometry, order, connectivity, and generation into a more unified framework. In contrast, graph neural network-based methods explicitly represent node relationships, edge constraints, and multi-scale connectivity patterns through graph structure. Hybrid methods further combine segmentation, sequence modeling, graph reasoning, and redundant supervision to improve robustness in complex road scenes. The main challenge is no longer limited to recovering local geometric shapes, but increasingly lies in maintaining structural consistency when map elements, topological relations, temporal information, and prior knowledge are considered together. Beyond these academic methods, industrial practice has also developed end-to-end map generation systems for large-scale urban deployment, aiming to improve automation, reduce production costs, and shorten update latency. Related studies also cover online high-definition map construction and pseudo-label learning under weak or missing annotations. Meanwhile, large models and vision-language models have gradually entered this line of research. Their potential in map generation has attracted increasing attention; however, their practical use is still constrained by data quality, structural priors, engineering controllability, and the requirements of real production environments. Conclusions and Prospects: Overall, end-to-end autonomous driving map generation is reshaping the conventional map construction paradigm and has shown clear advantages in process simplification, timeliness, and structured prediction. With the introduction of large models, vision-language models, and agent-based systems, map generation may gradually move beyond geometric construction toward a stage that also involves semantic understanding and human–machine interaction. Further progress in this direction will depend not only on model capability but also on the establishment of stable data pipelines, quality control mechanisms, and closed-loop engineering workflows, thereby supporting the large-scale and stable deployment of autonomous driving systems.
自动驾驶地图 / 端到端地图生成 / 工业级方案 / 视觉语言模型
autonomous driving map / end-to-end map generation / industrial-grade solution / vision language model
| 1 |
国家发展改革委, 中央网信办, 科技部, 等. 关于印发《智能汽车创新发展战略》的通知[EB/OL]. (2020-02-10)[2025-09-27]. https://www.gov.cn/zhengce/zhengceku/2020-02/24/content_5482655.htm
National Development and Reform Commission, Office of the Central Cyberspace Affairs Commission, Ministry of Science and Technology, et al. Notice on issuing the "smart vehicle innovation and development strategy"[EB/OL]. (2020-02-10)[2025-09-27]. https://www.gov.cn/zhengce/zhengceku/2020-02/24/content_5482655.htm. (in Chinese)
|
| 2 |
XIA D G, ZHANG W M, LIU X Y, et al. DuMapNet: An end-to-end vectorization system for city-scale lane-level map generation[C]//Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Barcelona, Spain: ACM, 2024: 6015–6024.
|
| 3 |
|
| 4 |
刘经南, 詹骄, 郭迟, 等. 智能高精地图数据逻辑结构与关键技术[J]. 测绘学报, 2019, 48(8): 939- 953.
|
| 5 |
杨蒙蒙, 江昆, 温拓朴, 等. 自动驾驶高精度地图众源更新技术现状与挑战[J]. 中国公路学报, 2023, 36(5): 244- 259.
|
| 6 |
张焱杰, 黄炜, 刘信陶, 等. 自动驾驶高精地图信息交互方法[J]. 武汉大学学报(信息科学版), 2024, 49(4): 662- 671.
|
| 7 |
王舒曼, 应申, 蒋跃文, 等. 智能驾驶场景中高精地图动静态数据关联方法[J]. 武汉大学学报(信息科学版), 2024, 49(4): 640- 650.
|
| 8 |
陈会仙, 章炜, 杨蒙蒙, 等. 面向无人驾驶的高精地图公开应用趋势研究[J]. 武汉大学学报(信息科学版), 2024, 49(4): 537- 545.
|
| 9 |
|
| 10 |
WIJAYA B, JIANG K, YANG M M, et al. High definition map mapping and update: A general overview and future directions [EB/OL]. (2024-09-15)[2025-09-27]. https://arxiv.org/abs/2409.09726.
|
| 11 |
应申, 蒋跃文, 顾江岩, 等. 面向自动驾驶的高精地图模型及关键技术[J]. 武汉大学学报(信息科学版), 2024, 49(4): 506- 515.
|
| 12 |
|
| 13 |
中国汽车工程学会. 智能网联汽车自动驾驶地图数据增量更新第1部分: 通用要求: T/CSAE 366.1—2024[S]. 北京: 中国汽车工程学会, 2024.
China Society of Automotive Engineers. Intelligent and connected vehicle—Incremental updating of automated driving map data—Part 1: General requirements: T/CSAE 366.1—2024[S]. Beijing: China Society of Automotive Engineers, 2024. (in Chinese)
|
| 14 |
自然资源部. 自然资源部关于加强智能网联汽车有关测绘地理信息安全管理的通知[EB/OL]. (2024-07-26)[2025-09-27]. https://www.gov.cn/zhengce/zhengceku/202407/content_6965227.htm.
Ministry of Natural Resources. Notice of the ministry of natural resources on strengthening the management of surveying and mapping geographic information security related to intelligent and connected vehicles[EB/OL]. (2024-07-26)[2025-09-27]. https://www.gov.cn/zhengce/zhengceku/202407/content_6965227.htm. (in Chinese)
|
| 15 |
ISO. Intelligent transport systems—Cooperative systems—State of art of Local Dynamic Map concepts: ISO/TR 17424: 2015[R]. Geneva: ISO, 2015.
|
| 16 |
HERRTWICH R. The evolution of the HERE HD live map at Daimler[EB/OL]. (2018-02-27)[2025-09-27]. https://www.here.com/learn/blog/the-evolution-of-the-hd-live-map.
|
| 17 |
MA C J. Baidu Apollo HD map[EB/OL]. (2018-11-21)[2025-09-27]. https://ggim.un.org/unwgic/presentations/2.2_Ma_Changjie.pdf.
|
| 18 |
四维图新. 四维图新 - 轻量级高精地图数据[EB/OL]. (2025)[2025-09-27]. https://www.seewayai.com/map-products.
NAVINFO. NavInfo - Lightweight HD map data[EB/OL]. (2025)[2025-09-27]. https://www.seewayai.com/map-products. (in Chinese)
|
| 19 |
LI Q, WANG Y, WANG Y L, et al. HDMapNet: An online HD map construction and evaluation framework[C]//2022 International Conference on Robotics and Automation. Philadelphia, USA: IEEE, 2022: 4628–4634.
|
| 20 |
LIU Y C, YUAN T Y, WANG Y, et al. VectorMapNet: End-to-end vectorized HD map learning[C]//Proceedings of the 40th International Conference on Machine Learning. Honolulu, USA: PMLR, 2023: 22352–22369.
|
| 21 |
LIAO B C, CHEN S Y, WANG X G, et al. MapTR: Structured modeling and learning for online vectorized HD map construction[C]//The Eleventh International Conference on Learning Representations. Kigali, Rwanda: OpenReview. net, 2023.
|
| 22 |
|
| 23 |
|
| 24 |
|
| 25 |
|
| 26 |
|
| 27 |
|
| 28 |
SANDSTRÖM E, LI Y, VAN GOOL L, et al. Point-SLAM: Dense neural point cloud-based SLAM[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023: 18387–18398.
|
| 29 |
|
| 30 |
ZHANG J, SINGH S. LOAM: Lidar odometry and mapping in real-time[C]//Robotics: Science and Systems X. Berkeley, USA: Robotics: Science and Systems Foundation, 2014: 1–9.
|
| 31 |
TSOTSOS K, CHIUSO A, SOATTO S. Robust inference for visual-inertial sensor fusion[C]//2015 IEEE International Conference on Robotics and Automation. Seattle, USA: IEEE, 2015: 5203–5210.
|
| 32 |
ZHAO S B, ZHANG H R, WANG P, et al. Super odometry: IMU-centric LiDAR-visual-inertial estimator for challenging environments[C]//2021 IEEE/RSJ International Conference on Intelligent Robots and Systems. Prague, Czech Republic: IEEE, 2021: 8729–8736.
|
| 33 |
|
| 34 |
SHAO W Z, VIJAYARANGAN S, LI C, et al. Stereo visual inertial LiDAR simultaneous localization and mapping[C]//2019 IEEE/RSJ International Conference on Intelligent Robots and Systems. Piscataway, NJ: IEEE, 2019: 370–377.
|
| 35 |
|
| 36 |
SEGAL A, HÄEHNEL D, THRUN S. Generalized-ICP[C]// Robotics: Science and Systems V. Seattle, USA: The MIT Press, 2009: 435.
|
| 37 |
FAN M, YAO Y, ZHANG J P, et al. Neural HD map generation from multiple vectorized tiles locally produced by autonomous vehicles[C]//5th China Conference on Spatial Data and Intelligence. Nanjing, China: Springer, 2024: 307–318.
|
| 38 |
戴激光, 王杨, 杜阳, 等. 光学遥感影像道路提取的方法综述[J]. 遥感学报, 2020, 24(7): 804- 823.
|
| 39 |
YANG J Z, YE X Q, WU B, et al. DuARE: Automatic road extraction with aerial images and trajectory data at Baidu maps [C]//Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Washington, USA: ACM, 2022: 4321–4331.
|
| 40 |
ZOU M, KAGEYAMA Y. Road curb detection based on a deep learning framework[C]//2023 IEEE 13th Annual Computing and Communication Workshop and Conference. Las Vegas, USA: IEEE, 2023: 259–262.
|
| 41 |
|
| 42 |
|
| 43 |
|
| 44 |
|
| 45 |
|
| 46 |
|
| 47 |
李必军, 郭圆, 周剑, 等. 智能驾驶高精地图发展与展望[J]. 武汉大学学报(信息科学版), 2024, 49(4): 491- 505.
|
| 48 |
LIU R J, YUAN Z J, LIU T, et al. End-to-end lane shape prediction with transformers[C]//Proceedings of the IEEE Winter Conference on Applications of Computer Vision. Waikoloa, USA: IEEE, 2021: 3693–3701.
|
| 49 |
CHEN L, SIMA C H, LI Y, et al. PersFormer: 3D lane detection via perspective transformer and the OpenLane benchmark[C]// 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022: 550–567.
|
| 50 |
CHEN S Y, CHENG T H, WANG X G, et al. Efficient and robust 2D-to-BEV representation learning via geometry-guided kernel transformer[EB/OL]. (2022-06-09)[2025-09-27]. https://arxiv.org/abs/2206.04584.
|
| 51 |
ZHOU B, KRÄHENBÜHL P. Cross-view transformers for real-time map-view semantic segmentation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans, USA: IEEE, 2022: 13750–13759.
|
| 52 |
|
| 53 |
LIU X L, WANG S, LI W T, et al. MGMap: Mask-guided learning for online vectorized HD map construction[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024: 14812–14821.
|
| 54 |
ZHANG X Y, LIU G W, LIU Z H, et al. Enhancing vectorized map perception with historical rasterized maps[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 422–439.
|
| 55 |
HAO X S, LI R K, ZHANG H, et al. MapDistill: Boosting efficient camera-based HD map construction via camera-LiDAR fusion model distillation[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 166–183.
|
| 56 |
ZHANG G J, LIN J H, WU S, et al. Online map vectorization for autonomous driving: A rasterization perspective[C]//Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023: 1382.
|
| 57 |
DING W J, QIAO L M, QIU X, et al. PivotNet: Vectorized pivot learning for end-to-end HD map construction[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023: 3649–3659.
|
| 58 |
YUAN T Y, LIU Y C, WANG Y, et al. StreamMapNet: Streaming mapping network for vectorized online HD map construction[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa, USA: IEEE, 2024: 7341–7350.
|
| 59 |
ZHANG Z X, ZHANG Y Y, DING X H, et al. Online vectorized HD map construction using geometry[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 73–90.
|
| 60 |
QIAO L M, DING W J, QIU X, et al. End-to-end vectorized HD-map construction with piecewise Bézier curve [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023: 13218–13228.
|
| 61 |
ZHOU Y, ZHANG H, YU J Q, et al. HIMap: Hybrid representation learning for end-to-end vectorized HD map construction[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024: 15396-15406.
|
| 62 |
|
| 63 |
LIU Z H, ZHANG X Y, LIU G W, et al. Leveraging enhanced queries of point sets for vectorized map construction[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 461–477.
|
| 64 |
PENG N, ZHOU X, WANG M M, et al. PrevPredMap: Exploring temporal modeling with previous predictions for online vectorized HD map construction[C]//2025 IEEE/CVF Winter Conference on Applications of Computer Vision. Tucson, USA: IEEE, 2025: 8134–8143.
|
| 65 |
YU J Y, ZHANG Z Z, XIA S F, et al. ScalableMap: Scalable map learning for online long-range vectorized HD map construction[C]//Proceedings of the 7th Conference on Robot Learning. Atlanta, USA: PMLR, 2023: 2429–2443.
|
| 66 |
XIONG X, LIU Y C, YUAN T Y, et al. Neural map prior for autonomous driving[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023: 17535–17544.
|
| 67 |
DONG H, GU W H, ZHANG X J, et al. SuperFusion: Multilevel LiDAR-camera fusion for long-range HD map generation [C]// 2024 IEEE International Conference on Robotics and Automation. Yokohama, Japan: IEEE, 2024: 9056–9062.
|
| 68 |
CHEN J C, WU Y F, TAN J Q, et al. MapTracker: Tracking with strided memory fusion for consistent vector HD mapping[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 90–107.
|
| 69 |
ZHANG D P, CHEN D Y, ZHI P, et al. MapExpert: Online HD map construction with simple and efficient sparse map element expert[C]//Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, USA: AAAI Press, 2025: 14745–14753.
|
| 70 |
|
| 71 |
LÖWENS C, FUNKE T, XIE J C, et al. PseudoMapTrainer: Learning online mapping without HD maps[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Seattle, USA: IEEE, 2025: 5263–5272.
|
| 72 |
XU Z H, LIU Y X, SUN Y X, et al. CenterLineDet: CenterLine graph detection for road lanes with vehicle-mounted sensors by transformer for HD map generation[C]//2023 IEEE International Conference on Robotics and Automation. London, UK: IEEE, 2023: 3553–3559.
|
| 73 |
HAN Y H, YU K, LI Z W. Continuity preserving online CenterLine graph learning[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 342–359.
|
| 74 |
CAN Y B, LINIGER A, PAUDEL D P, et al. Structured bird's-eye-view traffic scene understanding from onboard images [C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal, Canada: IEEE, 2021: 15641–15650.
|
| 75 |
|
| 76 |
LU J C, LI H Y, PENG R Y, et al. Translating images to road network: A non-autoregressive sequence-to-sequence approach[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023: 23–33.
|
| 77 |
FU Y P, LIAO W B, LIU X Y, et al. TopoLogic: An interpretable pipeline for lane topology reasoning on driving scenes [C]// Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024: 1970.
|
| 78 |
YANG Y M, LUO Y R, HE B K, et al. Topo2Seq: Enhanced topology reasoning via topology sequence learning[C]//Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, USA: AAAI Press, 2025: 9318–9326.
|
| 79 |
LI H, HUANG S F, XU L F, et al. RATopo: Improving lane topology reasoning via redundancy assignment[C]//Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025: 777–786.
|
| 80 |
PENG R Y, CAI X Y, XU H, et al. LaneGraph2Seq: Lane topology extraction with language model via vertex-edge encoding and connectivity enhancement[C]//Proceedings of the 38th AAAI Conference on Artificial Intelligence. Vancouver, Canada: AAAI Press, 2024: 4497–4505.
|
| 81 |
WU D M, CHANG J H, JIA F, et al. TopoMLP: A simple yet strong pipeline for driving topology reasoning[C]//The Twelfth International Conference on Learning Representations. Vienna, Austria: OpenReview. net, 2024.
|
| 82 |
LIAO B C, CHEN S Y, JIANG B, et al. Lane graph as path: Continuity-preserving path-wise modeling for online lane graph construction[C]//18th European Conference on Computer Vision. Milan, Italy: Springer, 2024: 334–351.
|
| 83 |
ZHU T Y, LENG J H, ZHONG J R, et al. LaneMapNet: Lane network recognization and HD map construction using curve region aware temporal bird's-eye-view perception[C]//2024 IEEE Intelligent Vehicles Symposium. Piscataway, NJ: IEEE, 2024: 2168–2175.
|
| 84 |
LI T Y, JIA P J, WANG B J, et al. LaneSegNet: Map learning with lane segment perception for autonomous driving[C]//The Twelfth International Conference on Learning Representations. Vienna, Austria: OpenReview. net, 2024.
|
| 85 |
JIA P J, WEN T P, LUO Z A, et al. LaneDAG: Automatic HD map topology generator based on geometry and attention fusion mechanism[C]//2024 IEEE Intelligent Vehicles Symposium. Piscataway, NJ: IEEE, 2024: 1015–1021.
|
| 86 |
MI L, ZHAO H, NASH C, et al. HDMapGen: A hierarchical graph generative model of high definition maps[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville, USA: IEEE, 2021: 4225–4234.
|
| 87 |
XIA D G, ZHANG W M, LIU X Y, et al. LDMapNet-U: An end-to-end system for city-scale lane-level map updating[C]// Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Toronto, Canada: ACM, 2025: 2693–2702.
|
| 88 |
|
| 89 |
LI T. MapNeXt: Revisiting training and scaling practices for online vectorized HD map construction[EB/OL]. (2024-01-14)[2025-09-27]. https://arxiv.org/abs/2401.07323.
|
| 90 |
|
| 91 |
OpenAI. GPT-4 technical report[EB/OL]. (2023-03-15)[2025-07-27]. https://arxiv.org/abs/2303.08774.
|
| 92 |
LIU H T, LI C Y, WU Q Y, et al. Visual instruction tuning[C]// Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023: 1516.
|
| 93 |
LIU H T, LI C Y, LI Y H, et al. Improved baselines with visual instruction tuning[C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024: 26286–26296.
|
| 94 |
LIU H T, LI C Y, LI Y H, et al. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge[EB/OL]. (2024-01-30)[2025-09-27]. https://llava-vl.github.io/blog/2024-01-30-llava-next/.
|
| 95 |
CAO X, ZHOU T, MA Y S, et al. MAPLM: A real-world large-scale vision-language benchmark for map and traffic scene understanding[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024: 21819–21830.
|
| 96 |
|
| 97 |
CHOUDHARY T, DEWANGAN V, CHANDHOK S, et al. Talk2BEV: Language-enhanced bird's-eye view maps for autonomous driving[C]//2024 IEEE International Conference on Robotics and Automation. Yokohama, Japan: IEEE, 2024: 16345–16352.
|
| 98 |
XU Q Y, CHEN S H, CHEN G, et al. ChatBEV: A visual language model that understands BEV maps[EB/OL]. (2025-03-18)[2025-09-27]. https://arxiv.org/abs/2503.13938.
|
/
| 〈 |
|
〉 |