基于深度强化学习的多目标柔性作业车间调度方法

A multi-objective flexible job shop scheduling method based on deep reinforcement learning

  • 摘要: 针对多目标柔性作业车间调度中最大完工时间与系统能耗难以协同优化的问题,提出一种基于图神经网络与条件化多目标近端策略优化算法(conditional multi-objective proximal policy optimization, CMO-PPO)的端到端调度方法。首先,将调度过程建模为多目标马尔可夫决策过程,以向量形式刻画完工时间与总能耗;其次,将调度系统表示为包含工序节点与机器节点的异构图结构,并利用图神经网络提取结构化状态特征;在此基础上,构建向量奖励与向量价值函数的多目标PPO框架,在优势层引入延迟标量化融合机制以实现稳定更新,同时通过偏好权重条件化策略提升模型在不同目标权衡下的泛化能力。实验结果表明,所提方法与多目标遗传算法(non-dominated sorting genetic algorithm II,NSGA-II)、快速单形体增长算法(fast N-simplex growth algorithm,FNSGA)和NRainbow方法相比,在多数算例上能够获得质量较高且覆盖范围更广的帕累托解集,表现出较好的多目标优化能力。

     

    Abstract: To address the challenge of jointly optimizing makespan and system energy consumption in multi-objective flexible job shop scheduling problems, an end-to-end scheduling method based on graph neural networks and conditional multi-objective proximal policy optimization (CMO-PPO) is proposed. Firstly, the scheduling process is formulated as a multi-objective Markov decision process, in which makespan and total energy consumption are represented as vectors. Subsequently, the scheduling system is modeled as a heterogeneous graph composed of operation nodes and machine nodes, and structured state features are encoded by a graph neural network. On this basis, a multi-objective PPO framework with vector rewards and vector value functions is constructed. A delayed scalarization fusion mechanism is introduced at the advantage estimation level to ensure stable policy updates, while preference-conditioned policies are employed to enhance the generalization capability of the model under different objective trade-offs. Experimental results demonstrate that, compared with NSGA-II, FNSGA, and NRainbow, the proposed method is capable of obtaining Pareto solution sets with relatively high quality and broader coverage in most benchmark instances, thereby exhibiting good multi-objective optimization performance.

     

/

返回文章
返回