基于元强化学习的应急救援边缘云服务资源缓存

邵沛祖, 周楚然, 蒋思为, 林尚静, 单维峰, 吴伟民, 李继龙

清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (9) : 1890-1901.

PDF(1929 KB)
PDF(1929 KB)
清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (9) : 1890-1901. DOI: 10.16511/j.cnki.qhdxxb.2026.27.052
公共安全

基于元强化学习的应急救援边缘云服务资源缓存

作者信息 +

Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services

Author information +
文章历史 +

摘要

面向灾害应急场景下边缘云服务请求突发、新场景样本有限导致的资源预缓存决策困难问题,该文提出一种基于元强化学习(meta-reinforcement learning, Meta-RL)的服务资源预缓存方法。该方法将资源预缓存过程建模为Markov决策过程(Markov decision process, MDP),以近端策略优化算法(proximal policy optimization, PPO)为基础策略框架,并从服务保障能力与资源成本效益2个维度构建综合评价指标,通过多任务训练学习可迁移的策略初始化参数,实现新场景下缓存策略的少样本快速适应。仿真结果表明,基于元强化学习的近端策略优化算法(METAPPO)在场景变化后仍具有优异的收敛性能。该方法能够提升灾害应急边缘云服务资源预缓存的快速适应能力,为少样本条件下的应急边缘智能服务资源配置提供支撑。

Abstract

Objective: In disaster emergencies, edge cloud service resource pre-caching faces two major challenges. First, post-disaster service demands exhibit strong burstiness and temporal dynamics. Insufficient pre-cached resources may cause request queuing and prolonged response latency, whereas excessive pre-caching wastes computing resources and incurs additional operating costs. Second, because disaster events are rare, historical data are insufficient for newly affected regions, new service types, and newly deployed edge nodes. Conventional prediction and reinforcement learning methods therefore struggle to obtain effective pre-caching policies rapidly under few-sample conditions. To address these challenges, this paper proposes an edge cloud service resource pre-caching method based on meta-reinforcement learning (Meta-RL). Methods: Because user behavior data from real disaster scenarios are scarce, this paper analyzes service traffic in routine operational settings and trains the agent through interactions with daily service environments. The goal is to learn transferable workload evolution patterns and adapt quickly to emergency service modes. User request traffic is analyzed across Internet data centers (IDCs) and service types. The results show that traffic variations among IDCs exhibit correlations at monthly and intraday scales, indicating transferable common structures across regional workloads. Meanwhile, differences in request baselines, peak intensities, and fluctuation ranges reveal environmental heterogeneity. Across service types, different services share temporal trends but differ in request scale, peak duration, dispersion degree, and tail-load distribution. These observations indicate that daily service scenarios provide transferable experience, whereas data distribution differences hinder independent reinforcement learning training. Therefore, a Meta-RL framework is constructed to learn transferable policy initialization parameters through multitask training. Results: In the proposed framework, the edge cloud service resource pre-caching process is formulated as a Markov decision process. The agent determines the number of resources to pre-cache in the next time slot based on the current resource queue length and request waiting queue length. The reward function jointly constrains redundant cached resources and waiting requests, guiding the agent to balance service responsiveness and resource efficiency. To evaluate this trade-off, a comprehensive pooling indicator is designed from request hit capability and resource waste rate. The hit rate measures the ability of pre-cached resources to immediately satisfy user requests, whereas the waste rate characterizes the proportion of cached resources that remain unused. For policy optimization, proximal policy optimization (PPO) is adopted as the basic framework, and a meta proximal policy optimization (METAPPO) algorithm is developed. METAPPO performs task-specific adaptation through inner-loop updates and optimizes shared initial policy parameters via outer-loop updates, improving adaptability to new IDCs, service types, and emergency tasks. Conclusions: Simulation experiments compare METAPPO with PPO and other baseline methods across edge cloud service environments. The results show that METAPPO accelerates convergence across multiple scenarios. After a scenario change, METAPPO still exhibits favorable convergence performance, demonstrating that the meta-learning mechanism improves policy adaptation efficiency. Although some conventional methods may achieve competitive steady-state performance after sufficient training, they usually require more historical samples and longer training processes. Conversely, METAPPO leverages cross-task experience from daily service scenarios and rapidly generates effective resource pre-caching policies under few-sample conditions, supporting fast resource allocation adaptation in multiregion, multiservice, and newly emerging emergency edge cloud scenarios.

关键词

灾害应急 / 边缘云服务 / 资源缓存 / 元强化学习

Key words

disaster emergency / edge cloud service / resource caching / meta-reinforcement learning

引用本文

导出引用
邵沛祖, 周楚然, 蒋思为, . 基于元强化学习的应急救援边缘云服务资源缓存[J]. 清华大学学报(自然科学版). 2026, 66(9): 1890-1901 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.052
Peizu SHAO, Churan ZHOU, Siwei JIANG, et al. Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(9): 1890-1901 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.052
中图分类号: TP393.1   

参考文献

1
SANTI N, MITTON N. A resource management survey for mission-critical and time-critical applications in multiaccess edge computing[J]. ITU Journal on Future and Evolving Technologies, 2021, 2(2): 61- 80.
2
MAHBUB M, SHUBAIR R M. Contemporary advances in multi-access edge computing: A survey of fundamentals, architecture, technologies, deployment cases, security, challenges, and directions[J]. Journal of Network and Computer Applications, 2023, 219, 103726.
3
AOUEDI O, LE V A, PIAMRAT K, et al. Deep learning on network traffic prediction: Recent advances, analysis, and future directions[J]. ACM Computing Surveys, 2025, 57(6): 151.
4
WANG Y, CHEN P Y. Network traffic prediction based on transformer and temporal convolutional network[J]. PLOS One, 2025, 20(4): e0320368.
5
宁兆龙, 张凯源, 王小洁, 等. 基于多智能体元强化学习的车联网协同服务缓存和计算卸载[J]. 通信学报, 2021, 42(6): 118- 130.
NING Z L, ZHANG K Y, WANG X J, et al. Cooperative service caching and peer offloading in Internet of vehicles based on multi-agent meta-reinforcement learning[J]. Journal on Communications, 2021, 42(6): 118- 130.
6
林鹏, 王俊, 刘艳, 等. 移动网络中基于元强化学习的多维性能自适应内容缓存策略[J]. 电子与信息学报, 2025, 47(8): 2598- 2607.
LIN P, WANG J, LIU Y, et al. Multi-dimensional performance adaptive content caching in mobile networks based on meta reinforcement learning[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2598- 2607.
7
WEI Z C, ZHAO Y, LYU Z W, et al. Cooperative caching algorithm for mobile edge networks based on multi-agent meta reinforcement learning[J]. Computer Networks, 2024, 242, 110247.
8
张依琳, 梁玉珠, 尹沐君, 等. 移动边缘计算中计算卸载方案研究综述[J]. 计算机学报, 2021, 44(12): 2406- 2430.
ZHANG Y L, LIANG Y Z, YIN M J, et al. Survey on the methods of computation offloading in mobile edge computing[J]. Chinese Journal of Computers, 2021, 44(12): 2406- 2430.
9
唐朝刚, 李召, 肖硕, 等. 一种面向车载边缘计算基于服务缓存的任务协同卸载算法[J]. 计算机学报, 2025, 48(4): 864- 876.
TANG C G, LI Z, XIAO S, et al. Service caching based collaborative task offloading algorithm in vehicular edge computing[J]. Chinese Journal of Computers, 2025, 48(4): 864- 876.
10
陈清林, 邝祝芳. 基于DDPG的边缘计算任务卸载和服务缓存算法[J]. 计算机工程, 2021, 47(10): 26- 33.
CHEN Q L, KUANG Z F. Task offloading and service caching algorithm for edge computing based on DDPG in edge computing[J]. Computer Engineering, 2021, 47(10): 26- 33.
11
KE H C, WANG H, YANG K, et al. Service caching decision-making policy for mobile edge computing using deep reinforcement learning[J]. IET Communications, 2023, 17(3): 362- 376.
12
XIE M D, YE J F, ZHANG G P, et al. Deep reinforcement learning-based computation offloading and distributed edge service caching for mobile edge computing[J]. Computer Networks, 2024, 250, 110564.
13
TANG C G, DING Y, XIAO S, et al. Collaborative service caching, task offloading, and resource allocation in caching-assisted mobile edge computing[J]. IEEE Transactions on Services Computing, 2025, 18(4): 1966- 1981.
14
PENG Z Y, QIU Y, WANG G C. Model caching and application offloading for mobile edge intelligence network with learning-and-optimization approach[J]. IEEE Transactions on Services Computing, 2025, 18(5): 3022- 3037.
15
涂静正, 温晓婧, 陈彩莲, 等. 基于边缘计算的工业视频网络智能感知: 挑战与进展[J]. 自动化学报, 2025, 51(8): 1715- 1738.
TU J Z, WEN X J, CHEN C L, et al. Edge computing based intelligent perception for industrial video network: Challenge and progress[J]. Acta Automatica Sinica, 2025, 51(8): 1715- 1738.

基金

复杂零部件智能检测与识别湖北省工程研究中心开放课题(IDICP-KF-2024-25)
北京市自然科学基金-海淀原始创新联合基金(L232001)

版权

版权所有,未经授权,不得转载。
PDF(1929 KB)

Accesses

Citation

Detail

段落导航
相关文章

/