线性规划最大值优化

2条回答

网友

1楼 · 编辑于 2024-09-30 23:39:21

下面是关于我使用最小成本流的一些想法

我们用一个有向图来模拟这个问题，图中有4层，每一层与下一层完全相连

节点

第一层：一个节点s，它将是我们的源
第二层：每个学生一个节点
第三层：每个主题一个节点
第四层：OA单节点t将成为我们的排水管

边缘容量

第一->；第二：所有边的容量为1
第二->；第三：所有边的容量为1
第三->；第四：所有边缘都具有与必须分配给该科目的学生数量相对应的容量

边缘成本

第一->；第二：所有边的成本为0
第二->；第三：记住，这一层中的边缘将学生与主题联系起来。这些课程的费用将与学生在该科目上的加权分数成反比。 cost = -subject_weight*student_subject_score
第三->；第四：所有边的成本为0

然后我们要求从s到t的流量等于我们必须选择的学生数量

为什么这样做？

通过将第三层和第四层之间的所有边作为赋值，最小成本流问题的解决方案将与您的问题的解决方案相对应

每个学生最多可以选择一个主题，因为对应的节点只有一个传入边缘

每门学科都有确切的所需学生人数，因为我们必须为这门学科选择的学生人数对应着离校能力，我们必须充分利用这些边缘的能力，否则我们无法满足流动需求

MCF问题的最小解对应于问题的最大解，因为成本对应于它们给出的值

当您要求用python提供解决方案时，我用ortools实现了最小成本流问题。在我的colab笔记本上找到一个解决方案不到一秒钟。“长时间”需要的是提取溶液。但包括设置和解决方案提取，对于整个100000名学生的问题，我的运行时间仍然不到20秒

代码

# imports
from ortools.graph import pywrapgraph
import numpy as np
import pandas as pd
import time
t_start = time.time()

# setting given problem parameters
num_students = 100000
subjects = ['MATH', 'ENGLISH', 'COMPUTERS', 'HISTORY','PHYSICS']
num_subjects = len(subjects)
demand = [4000, 3000, 2000, 750, 250]
weight = [1.9,1.7, 1.5, 1.3, 1.1]

# generating student scores
student_scores_raw = np.random.randint(101, size=(num_students, num_subjects))

# setting up graph nodes
source_nodes = [0]
student_nodes = list(range(1, num_students+1))
subject_nodes = list(range(num_students+1, num_subjects+num_students+1))
drain_nodes = [num_students+num_subjects+1]

# setting up the min cost flow edges
start_nodes = [int(c) for c in (source_nodes*num_students + [i for i in student_nodes for _ in subject_nodes] + subject_nodes)]
end_nodes   = [int(c) for c in (student_nodes + subject_nodes*num_students + drain_nodes*num_subjects)]
capacities  = [int(c) for c in ([1]*num_students + [1]*num_students*num_subjects + demand)]
unit_costs  = [int(c) for c in ([0.]*num_students + list((-student_scores_raw*np.array(weight)*10).flatten()) + [0.]*num_subjects)]
assert len(start_nodes) == len(end_nodes) == len(capacities) == len(unit_costs)

# setting up the min cost flow demands
supplies = [sum(demand)] + [0]*(num_students + num_subjects) + [-sum(demand)]

# initialize the min cost flow problem instance
min_cost_flow = pywrapgraph.SimpleMinCostFlow()
for i in range(0, len(start_nodes)):
  min_cost_flow.AddArcWithCapacityAndUnitCost(start_nodes[i], end_nodes[i], capacities[i], unit_costs[i])
for i in range(0, len(supplies)):
  min_cost_flow.SetNodeSupply(i, supplies[i])

# solve the problem
t_solver_start = time.time()
if min_cost_flow.Solve() == min_cost_flow.OPTIMAL:
  print('Best Value:', -min_cost_flow.OptimalCost()/10)
  print('Solver time:', str(time.time()-t_solver_start)+'s')
  print('Total Runtime until solution:', str(time.time()-t_start)+'s')
  
  #extracting the solution
  solution = []
  for i in range(min_cost_flow.NumArcs()):
    if min_cost_flow.Flow(i) > 0 and min_cost_flow.Tail(i) in student_nodes:
      student_id = min_cost_flow.Tail(i)-1
      subject_id = min_cost_flow.Head(i)-1-num_students
      solution.append([student_id, subjects[subject_id], student_scores_raw[student_id, subject_id]])
  assert(len(solution) == sum(demand))

  solution = pd.DataFrame(solution, columns = ['student_id', 'subject', 'score'])
  print(solution.head())
else:
  print('There was an issue with the min cost flow input.')

print('Total Runtime:', str(time.time()-t_start)+'s')

在上面的代码中，用下面的列表理解（也不是每次迭代都使用列表查找）替换解决方案提取的for循环，可以显著提高运行时。但是出于可读性的原因，我也将这个旧的解决方案留在这里。这是新的：

  solution = [[min_cost_flow.Tail(i)-1, 
               subjects[min_cost_flow.Head(i)-1-num_students], 
               student_scores_raw[min_cost_flow.Tail(i)-1, min_cost_flow.Head(i)-1-num_students]
               ]
              for i in range(min_cost_flow.NumArcs())
              if (min_cost_flow.Flow(i) > 0 and 
                  min_cost_flow.Tail(i) <= num_students and 
                  min_cost_flow.Tail(i)>0)
              ]

下面的输出给出了新的更快实现的运行时

输出

Best Value: 1675250.7
Solver time: 0.542395830154419s
Total Runtime until solution: 1.4248979091644287s
   student_id    subject  score
0           3    ENGLISH     99
1           5       MATH     98
2          17  COMPUTERS    100
3          22  COMPUTERS    100
4          33    ENGLISH    100
Total Runtime: 1.752336025238037s

请指出我可能犯的任何错误

我希望这有帮助

网友

2楼 · 编辑于 2024-09-30 23:39:21

我想你已经接近了。这是一个相当标准的整数线性规划（ILP）分配问题。这会有点慢，因为问题的结构

你没有在你的帖子中说设置的故障是什么&；解决时间很短。我看到你在读一个文件并使用熊猫。我认为熊猫在优化问题上会很快变得笨重，但这只是个人喜好

我使用cbc解算器在pyomo中对您的问题进行了编码，我很确定这与pulp用于比较的解算器相同。（见下文）。我认为你有两个约束和一个双索引二进制决策变量

如果我把它减少到10万名学生（没有松弛…只是一对一配对），它会在14秒内解决，以便进行比较。我的设置是一个5年的iMac，带有大量ram

与池中的10万名学生一起运行，在调用解算器之前，它将在大约25分钟内以10秒的“设置”时间解算。所以我不确定为什么你的编码要花2小时。如果你能分解你的求解时间，那会有帮助。其余的应该是微不足道的。我没有在输出中做太多的改动，但OBJ函数值980K似乎是合理的

其他想法：

如果您可以正确配置解算器选项，并将mip间距设置为0.05左右，那么如果您可以接受稍微非最佳的解决方案，应该可以加快速度。我只在像古罗比这样的付费解算器的解算器选项上有过不错的运气。我没有太多的运气与使用免费解决方案，YMMV

import pyomo.environ as pyo
from random import randint
from time import time

# start setup clock
tic = time()
# exam types
subjects = ['Math', 'English', 'Computers', 'History', 'Physics']

# make set of students...
num_students = 100_000
students = [f'student_{s}' for s in range(num_students)]

# make 100K fake scores in "flat" format
student_scores = { (student, subj) : randint(0,100) 
                        for student in students
                        for subj in subjects}

assignments = { 'Math': 4000, 'English': 3000, 'Computers': 2000, 'History': 750, 'Physics': 250}
weights = {'Math': 1.9, 'English': 1.7, 'Computers': 1.5, 'History': 1.3, 'Physics': 1.1}

# Set up model
m = pyo.ConcreteModel('exam assignments')

# Sets
m.subjects = pyo.Set(initialize=subjects)
m.students = pyo.Set(initialize=students)

# Parameters
m.assignments = pyo.Param(m.subjects, initialize=assignments)
m.weights =     pyo.Param(m.subjects, initialize=weights)
m.scores =      pyo.Param(m.students, m.subjects, initialize=student_scores)

# Variables
m.x = pyo.Var(m.students, m.subjects, within=pyo.Binary)  # binary selection of pairing student to test

# Objective
m.OBJ = pyo.Objective(expr=sum(m.scores[student, subject] * m.x[student, subject] 
                for student in m.students
                for subject in m.subjects), sense=pyo.maximize)

### Constraints ###
# fill all assignments
def fill_assignments(m, subject):
    return sum(m.x[student, subject] for student in m.students) == assignments[subject]
m.C1 = pyo.Constraint(m.subjects, rule=fill_assignments)

# use each student at most 1 time
def limit_student(m, student):
    return sum(m.x[student, subject] for subject in m.subjects) <= 1
m.C2 = pyo.Constraint(m.students, rule=limit_student)

toc = time()
print (f'setup time: {toc-tic:0.3f}')
tic = toc

# solve it..
solver = pyo.SolverFactory('cbc')
solution = solver.solve(m)
print(solution)
toc = time()
print (f'solve time: {toc-tic:0.3f}')

输出

setup time: 10.835

Problem: 
- Name: unknown
  Lower bound: -989790.0
  Upper bound: -989790.0
  Number of objectives: 1
  Number of constraints: 100005
  Number of variables: 500000
  Number of binary variables: 500000
  Number of integer variables: 500000
  Number of nonzeros: 495094
  Sense: maximize
Solver: 
- Status: ok
  User time: -1.0
  System time: 1521.55
  Wallclock time: 1533.36
  Termination condition: optimal
  Termination message: Model was solved to optimality (subject to tolerances), and an optimal solution is available.
  Statistics: 
    Branch and bound: 
      Number of bounded subproblems: 0
      Number of created subproblems: 0
    Black box: 
      Number of iterations: 0
  Error rc: 0
  Time: 1533.8383190631866
Solution: 
- number of solutions: 0
  number of solutions displayed: 0

solve time: 1550.528

其他想法：

输出

相关问题更多 >

编程相关推荐

热门问题

热门文章