School of Information Technology and Engineering
Permanent URI for this collection
Browse
Browsing School of Information Technology and Engineering by Subject "Adaptive Fairness"
Now showing 1 - 1 of 1
Results Per Page
Sort Options
Item Adaptive Welfare Weighting for Enhanced Fairness Integration in Cooperative Multi-Agent Proximal Policy Optimization(Addis Ababa University, 2026-02) Bizuhan Abate Tayu; Beakal GizachewCooperative Multi-Agent Reinforcement Learning (MARL) primarily focuses on maximizing the cumulative reward to achieve high system efficiency. However, this singular focus often leads to a systemic situation where early success allows a subset of dominant agents to monopolize rewards which reflect Matthew Effect. These results long-term disparities in resource allocation among agents and the marginalization of agents. Matthew effect or rich-get-richer dynamic is particularly destructive in environments requiring sustained coordination, as it results in unfair reward distribution among agents, leading agent starvation. To address this, we propose the Weight Adaptive Fairness Integrated Proximal Policy Optimization (AW-FIPPO) algorithm. AW-FIPPO advance static fairness regularizers by introducing a dedicated Fairness Critic Network (F). This network processes the state of the environment and agents reward distribution to generate contextaware weight vectors (wt). These weights are then used to dynamically modulate a hybrid objective function that blends the standard PPO clipped surrogate loss with a concave, logarithmic utility transformation, enabling the system to automatically transitioning from utilitarian maximization to egalitarian equity based on the learned, state dependent weight. Empirical results across two distinct cooperative multi-agent environments demonstrate that the adaptive weighting mechanism is the superior engine for equitable outcomes and enhanced the welfare score and minimal agent reward while maintaining competitive total return against the baselines. The improvements are significantly larger in the Matthew Effect environment, where reward accumulation bias is inherent. In this setting, our method AW-FIPPO outperforms FAPPO by 122.1% gain in total score and a 84.3% gain in max score while reducing the Gini index and the Coefficient of Variation (CV) by 24.5% and 30.6%, respectively. Additionally, our attention-based variation, AT(AW-FIPPO) compared to baseline AT-FAPPO yields a 79.5% improvement in total score and a 40.1% gain in max score, alongside reductions of 27.5% in the Gini index and 29.8% in CV.