Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services

Peizu SHAO, Churan ZHOU, Siwei JIANG, Shangjing LIN, Weifeng SHAN, Weimin WU, Jilong LI

Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (9) : 1890-1901.

PDF(1929 KB)
PDF(1929 KB)
Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (9) : 1890-1901. DOI: 10.16511/j.cnki.qhdxxb.2026.27.052
Public Safety

Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services

Author information +
History +

Abstract

Objective: In disaster emergencies, edge cloud service resource pre-caching faces two major challenges. First, post-disaster service demands exhibit strong burstiness and temporal dynamics. Insufficient pre-cached resources may cause request queuing and prolonged response latency, whereas excessive pre-caching wastes computing resources and incurs additional operating costs. Second, because disaster events are rare, historical data are insufficient for newly affected regions, new service types, and newly deployed edge nodes. Conventional prediction and reinforcement learning methods therefore struggle to obtain effective pre-caching policies rapidly under few-sample conditions. To address these challenges, this paper proposes an edge cloud service resource pre-caching method based on meta-reinforcement learning (Meta-RL). Methods: Because user behavior data from real disaster scenarios are scarce, this paper analyzes service traffic in routine operational settings and trains the agent through interactions with daily service environments. The goal is to learn transferable workload evolution patterns and adapt quickly to emergency service modes. User request traffic is analyzed across Internet data centers (IDCs) and service types. The results show that traffic variations among IDCs exhibit correlations at monthly and intraday scales, indicating transferable common structures across regional workloads. Meanwhile, differences in request baselines, peak intensities, and fluctuation ranges reveal environmental heterogeneity. Across service types, different services share temporal trends but differ in request scale, peak duration, dispersion degree, and tail-load distribution. These observations indicate that daily service scenarios provide transferable experience, whereas data distribution differences hinder independent reinforcement learning training. Therefore, a Meta-RL framework is constructed to learn transferable policy initialization parameters through multitask training. Results: In the proposed framework, the edge cloud service resource pre-caching process is formulated as a Markov decision process. The agent determines the number of resources to pre-cache in the next time slot based on the current resource queue length and request waiting queue length. The reward function jointly constrains redundant cached resources and waiting requests, guiding the agent to balance service responsiveness and resource efficiency. To evaluate this trade-off, a comprehensive pooling indicator is designed from request hit capability and resource waste rate. The hit rate measures the ability of pre-cached resources to immediately satisfy user requests, whereas the waste rate characterizes the proportion of cached resources that remain unused. For policy optimization, proximal policy optimization (PPO) is adopted as the basic framework, and a meta proximal policy optimization (METAPPO) algorithm is developed. METAPPO performs task-specific adaptation through inner-loop updates and optimizes shared initial policy parameters via outer-loop updates, improving adaptability to new IDCs, service types, and emergency tasks. Conclusions: Simulation experiments compare METAPPO with PPO and other baseline methods across edge cloud service environments. The results show that METAPPO accelerates convergence across multiple scenarios. After a scenario change, METAPPO still exhibits favorable convergence performance, demonstrating that the meta-learning mechanism improves policy adaptation efficiency. Although some conventional methods may achieve competitive steady-state performance after sufficient training, they usually require more historical samples and longer training processes. Conversely, METAPPO leverages cross-task experience from daily service scenarios and rapidly generates effective resource pre-caching policies under few-sample conditions, supporting fast resource allocation adaptation in multiregion, multiservice, and newly emerging emergency edge cloud scenarios.

Key words

disaster emergency / edge cloud service / resource caching / meta-reinforcement learning

Cite this article

Download Citations
Peizu SHAO , Churan ZHOU , Siwei JIANG , et al . Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(9): 1890-1901 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.052

References

1
SANTI N, MITTON N. A resource management survey for mission-critical and time-critical applications in multiaccess edge computing[J]. ITU Journal on Future and Evolving Technologies, 2021, 2(2): 61- 80.
2
MAHBUB M, SHUBAIR R M. Contemporary advances in multi-access edge computing: A survey of fundamentals, architecture, technologies, deployment cases, security, challenges, and directions[J]. Journal of Network and Computer Applications, 2023, 219, 103726.
3
AOUEDI O, LE V A, PIAMRAT K, et al. Deep learning on network traffic prediction: Recent advances, analysis, and future directions[J]. ACM Computing Surveys, 2025, 57(6): 151.
4
WANG Y, CHEN P Y. Network traffic prediction based on transformer and temporal convolutional network[J]. PLOS One, 2025, 20(4): e0320368.
5
宁兆龙, 张凯源, 王小洁, 等. 基于多智能体元强化学习的车联网协同服务缓存和计算卸载[J]. 通信学报, 2021, 42(6): 118- 130.
NING Z L, ZHANG K Y, WANG X J, et al. Cooperative service caching and peer offloading in Internet of vehicles based on multi-agent meta-reinforcement learning[J]. Journal on Communications, 2021, 42(6): 118- 130.
6
林鹏, 王俊, 刘艳, 等. 移动网络中基于元强化学习的多维性能自适应内容缓存策略[J]. 电子与信息学报, 2025, 47(8): 2598- 2607.
LIN P, WANG J, LIU Y, et al. Multi-dimensional performance adaptive content caching in mobile networks based on meta reinforcement learning[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2598- 2607.
7
WEI Z C, ZHAO Y, LYU Z W, et al. Cooperative caching algorithm for mobile edge networks based on multi-agent meta reinforcement learning[J]. Computer Networks, 2024, 242, 110247.
8
张依琳, 梁玉珠, 尹沐君, 等. 移动边缘计算中计算卸载方案研究综述[J]. 计算机学报, 2021, 44(12): 2406- 2430.
ZHANG Y L, LIANG Y Z, YIN M J, et al. Survey on the methods of computation offloading in mobile edge computing[J]. Chinese Journal of Computers, 2021, 44(12): 2406- 2430.
9
唐朝刚, 李召, 肖硕, 等. 一种面向车载边缘计算基于服务缓存的任务协同卸载算法[J]. 计算机学报, 2025, 48(4): 864- 876.
TANG C G, LI Z, XIAO S, et al. Service caching based collaborative task offloading algorithm in vehicular edge computing[J]. Chinese Journal of Computers, 2025, 48(4): 864- 876.
10
陈清林, 邝祝芳. 基于DDPG的边缘计算任务卸载和服务缓存算法[J]. 计算机工程, 2021, 47(10): 26- 33.
CHEN Q L, KUANG Z F. Task offloading and service caching algorithm for edge computing based on DDPG in edge computing[J]. Computer Engineering, 2021, 47(10): 26- 33.
11
KE H C, WANG H, YANG K, et al. Service caching decision-making policy for mobile edge computing using deep reinforcement learning[J]. IET Communications, 2023, 17(3): 362- 376.
12
XIE M D, YE J F, ZHANG G P, et al. Deep reinforcement learning-based computation offloading and distributed edge service caching for mobile edge computing[J]. Computer Networks, 2024, 250, 110564.
13
TANG C G, DING Y, XIAO S, et al. Collaborative service caching, task offloading, and resource allocation in caching-assisted mobile edge computing[J]. IEEE Transactions on Services Computing, 2025, 18(4): 1966- 1981.
14
PENG Z Y, QIU Y, WANG G C. Model caching and application offloading for mobile edge intelligence network with learning-and-optimization approach[J]. IEEE Transactions on Services Computing, 2025, 18(5): 3022- 3037.
15
涂静正, 温晓婧, 陈彩莲, 等. 基于边缘计算的工业视频网络智能感知: 挑战与进展[J]. 自动化学报, 2025, 51(8): 1715- 1738.
TU J Z, WEN X J, CHEN C L, et al. Edge computing based intelligent perception for industrial video network: Challenge and progress[J]. Acta Automatica Sinica, 2025, 51(8): 1715- 1738.

RIGHTS & PERMISSIONS

All rights reserved. Unauthorized reproduction is prohibited.
PDF(1929 KB)

Accesses

Citation

Detail

Sections
Recommended

/