Multi-agent reinforcement learning for post-disaster waste management on dynamic road networks: A multi-agent proximal policy optimization approach with hard action masking
APPLIED SOFT COMPUTING, cilt.204, 2027 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 204
- Basım Tarihi: 2027
- Doi Numarası: 10.1016/j.asoc.2026.116438
- Dergi Adı: APPLIED SOFT COMPUTING
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, Compendex, INSPEC
- Kocaeli Üniversitesi Adresli: Evet
Özet
Post-disaster waste management poses a formidable combinatorial challenge that combines the classical Vehicle Routing Problem (VRP) with stochastic waste generation, dynamically deteriorating road networks, and multi-objective optimisation under severe uncertainty. A Multi-Agent Proximal Policy Optimisation (MAPPO) framework is proposed, operating under the Centralised Training, Decentralised Execution (CTDE) paradigm, to coordinate a heterogeneous fleet of waste-collection vehicles across earthquake-affected urban networks. The transportation infrastructure is modelled as a directed graph whose edge-health coefficients evolve via a com pound Poisson damage process with concurrent linear repair, while waste production at demolition sites follows a Log-Normal distribution with exponentially decaying mean. A hard action masking mechanism assigns zero selection probability to actions that are infeasible under the environment's feasibility information-travel on severely damaged roads, pickups at empty nodes, and capacity-violating loads-eliminating constraint violations within the simulator; its robustness to noisy state information is also quantified. MAPPO is evaluated against three online baselines (Nearest Neighbour, Clarke-Wright Savings, Genetic Algorithm), a single-agent PPO ab lation, and a static full-information Mixed-Integer Linear Programming reference, across four disaster scenarios of escalating severity, over 30 independent random seeds with 95% confidence intervals and paired significance tests. MAPPO attains the best multi-objective reward in every scenario (all p < 0.001). On the primary S2-Medium configuration it reduces total cost by 46% and carbon emissions by 51% relative to the best-performing baseline (the Genetic Algorithm), and by 69% and 73% respectively relative to Nearest Neighbour, while attaining the high est service level (1.6 & times; that of the nearest baseline). On the smallest and most severe tiers the Genetic Algorithm attains lower raw cost, so the comparison is reported as a trade-off. An action-masking ablation shows that hard masking is essential (a no-mask policy fails to learn), the CTDE factorisation outperforms a monolithic single-agent controller trained under the same budget, and the policy transfers zero-shot to damage intensities up to 4 & times; those seen in training. Reward-weight and fleet-capacity analyses further show that the objective weights act as interpretable levers over the cost-emission-service trade-off and that the low absolute service level reflects a structural capacity-demand asymmetry rather than a policy limitation. Controlled sensitivity experiments over fleet size, planning horizon, demand intensity and the road-passability threshold confirm that the framework scales gracefully and, in particular, is robust to the choice of the masking threshold.