面向动力煤分选过程的多智能体协同优化决策方法研究

    Study on a multi-agent collaborative optimization decision-making method for the thermal coal preparation process

    • 摘要: 为解决动力煤分选过程中脱粉粒度固定、“预先脱粉+重介质分选”工序独立调控、煤质波动下难以实现效益最优等问题,提出一种基于多智能体深度强化学习的协同优化决策方法,将原煤脱粉与重介质分选两个关键环节抽象为协作智能体,设计兼顾经济效益与产品质量的团队奖励函数,引入注意力机制优化策略网络,构建数字孪生仿真环境,依托集中式训练分布式执行框架完成智能体协同训练与仿真测试。为验证该策略方法的优越性,设置固定参数策略与单智能体独立控制策略开展对比分析。结果表明:在典型煤质波动工况下,该策略方法累计经济效益提升5.6%以上,精煤灰分合格率可达99.8%,分选密度波动均值为0.004 g/cm3;面对煤质突变的极端工况,仍可维持正向经济效益增长,在细粒级含量 > 50%的恶劣工况下经济效益仍可提升1.5%。该方法可打破选煤系统单环节局部优化局限,实现原煤脱粉与重介质分选跨工序协同调控,为选煤智能化由“感知—控制”迈向“认知—决策”提供理论与方法支撑。

       

      Abstract: To address the problems of a fixed desliming particle size, independent control of the “pre-deslim-ing + dense-medium separation” processes, and the difficulty in achieving optimal economic benefits under fluctuating coal quality during thermal coal preparation, a multi-agent deep reinforcement learning-based collaborative optimization decision-making method is proposed. The raw coal desliming and dense-medium separation processes are modeled as two cooperative agents, and a team reward function that considers both economic benefits and product quality is designed. An attention mechanism is introduced to optimize the policy network, and a digital twin-based simulation environment is developed. Based on the Centralized Training with Decentralized Execution framework, collaborative training and simulation testing of the agents are conducted. To verify the superiority of this strategy, comparative analyses were performed against a fixed-parameter strategy and an independent single-agent control strategy. The results show that under typical coal quality fluctuation conditions, this strategy increases cumulative economic benefits by more than 5.6%, achieves a clean coal ash content qualification rate of up to 99.8%, and maintains an average fluctuation in separation density of 0.004 g/cm3. Under extreme conditions involving abrupt changes in coal quality, the strategy still maintains positive economic growth. Even under severe conditions where the fine-particle content exceeds 50%, economic benefits can still be improved by 1.5%. The proposed method overcomes the limitations of local optimization at individual stages of the coal preparation system and enables cross-process collaborative control of raw coal desliming and dense-medium separation, providing theoretical and methodological support for the transformation of coal preparation intelligence from “perception–control” toward “cognition–decision-making.”

       

    /

    返回文章
    返回