Fairness-Aware Multi-Agent DRL for Sustainable Virtual Power Plants Operation Under Multi-Uncertainties面向多重不确定性的可持续虚拟电厂公平感知多智能体深度强化学习
英文 Abstract
The operation of virtual power plants (VPPs) is challenged by multi-source uncertainties in wholesale prices, PV output, and demand. However, existing studies mainly focus on market–operator interactions and have not sufficiently established fair and effective risk- and benefit-sharing between the VPP operator and consumers, which weakens operator incentives and consumers’ willingness to participate in demand response. To address this gap, we develop a fairness-aware bi-level coordination framework with a multi-agent deep reinforcement learning algorithm (MA-DSASAC). At the core, we embed a dynamic satisfaction module that models consumers’ bounded rationality and reference-dependent preferences, thereby preventing the ratchet effect. Combined with an economically constrained demand-response mechanism, it dynamically updates the satisfaction levels of both the VPP operator and consumers to incentivize proactive participation and enable fair risk sharing. For temporal modeling, we integrate a GRU and an attention mechanism to encode non-stationary price/PV/load sequences, ensuring stable convergence and improved generalization under multi-uncertainty. Using three years of real-world 5-min data, compared with industry TOU and flat-rate tariffs the proposed scheme increases operator revenue by 11–48% and reduces consumers’ electricity costs by 17–36% across both volatile and stable months; it further achieves peak-shaving/valley-filling of 14.9%/12.3% and 13.2%/10.9% in representative weeks, with Pareto improvements in 91.3% of operating periods. These results indicate stable and effective performance under multiple uncertainties, significantly enhancing VPP economic viability and load-response capability for the sustainable development of VPPs.
中文翻译
虚拟电厂运行面临批发电价、光伏出力和需求等多源不确定性的挑战。既有研究主要关注市场与运营商之间的交互,尚未充分建立虚拟电厂运营商与消费者之间公平有效的风险和收益分担机制,从而削弱运营商激励以及消费者参与需求响应的意愿。为弥补这一不足,本文构建兼顾公平性的双层协调框架,并提出多智能体深度强化学习算法 MA-DSASAC。其核心是引入动态满意度模块,对消费者的有限理性和参照依赖偏好进行建模,以避免棘轮效应;再结合具有经济约束的需求响应机制,动态更新运营商和消费者双方的满意度,激励主动参与并实现公平的风险分担。在时序建模方面,研究融合 GRU 与注意力机制,对非平稳的电价、光伏和负荷序列进行编码,从而在多重不确定性下实现稳定收敛并提高泛化能力。基于三年真实世界的五分钟粒度数据,与行业分时电价和平电价方案相比,所提方案在不同月份均提高了运营商收益并降低了消费者电费,同时改善了削峰填谷效果;绝大多数运行时段实现了帕累托改进。结果表明,该方法在多种不确定条件下具有稳定有效的表现,可增强虚拟电厂的经济可行性和负荷响应能力。
中文概述
本文针对虚拟电厂面临的电价、光伏出力和负荷不确定性,提出兼顾公平性的双层协调框架和多智能体深度强化学习算法。模型结合动态满意度、GRU 与注意力机制,在真实数据测试中改善运营商收益、用户电费及削峰填谷效果。