numpy返回一个3d数组内的索引

2条回答

网友

1楼 · 编辑于 2024-09-28 21:38:51

方法1

下面是一个使用^{}-

R,C = np.where((A[:,None,None] == B).any(-1))
out = np.split(C,np.flatnonzero(R[1:]>R[:-1])+1)

方法2

假设A和B都是正数，我们可以考虑它们来表示2D网格上的索引，这样{}可以被视为按行保存列索引。一旦与B对应的2D网格就位，我们只需要考虑与A相交的列。最后，我们得到这样一个2D网格中True值的索引，从而给出R和{}值。这应该更节省内存。在

因此，另一种方法应该是这样的-

^{pr2}$

样本运行-

In [43]: A
Out[43]: array([0, 1, 2, 3])

In [44]: B
Out[44]: 
array([[3, 2, 0],
       [0, 2, 1],
       [2, 3, 1],
       [3, 0, 1]])

In [45]: out
Out[45]: [array([0, 1, 3]), array([1, 2, 3]), array([0, 1, 2]), array([0, 2, 3])]

运行时测试

按100x放大数据集大小，下面是一个快速的运行时测试结果-

In [85]: def index_1din2d(A,B):
    ...:     R,C = np.where((A[:,None,None] == B).any(-1))
    ...:     out = np.split(C,np.flatnonzero(R[1:]>R[:-1])+1)
    ...:     return out
    ...: 
    ...: def index_1din2d_initbased(A,B):
    ...:     ncols = B.max()+1
    ...:     nrows = B.shape[0]
    ...:     mask = np.zeros((nrows,ncols),dtype=bool)
    ...:     mask[np.arange(nrows)[:,None],B] = 1
    ...:     mask[:,~np.in1d(np.arange(mask.shape[1]),A)] = 0
    ...:     R,C = np.where(mask.T)
    ...:     out = np.split(C,np.flatnonzero(R[1:]>R[:-1])+1)
    ...:     return out
    ...: 

In [86]: A = np.unique(np.random.randint(0,10000,(400)))
    ...: B = np.random.randint(0,10000,(400,300))
    ...: 

In [87]: %timeit [np.where((B == x).sum(axis = 1))[0] for x in A]
1 loop, best of 3: 161 ms per loop # @Psidom's soln

In [88]: %timeit index_1din2d(A,B)
10 loops, best of 3: 91.5 ms per loop

In [89]: %timeit index_1din2d_initbased(A,B)
10 loops, best of 3: 33.4 ms per loop

性能进一步提升！

或者，我们可以在第二种方法中以一种转置的方式创建2D网格。这个想法是为了避免R,C = np.where(mask.T)中的转置，这似乎是一个瓶颈。因此，第二种方法的修改版本和相关的运行时将如下所示-

In [135]: def index_1din2d_initbased_v2(A,B):
     ...:     nrows = B.max()+1
     ...:     ncols = B.shape[0]
     ...:     mask = np.zeros((nrows,ncols),dtype=bool)
     ...:     mask[B,np.arange(ncols)[:,None]] = 1
     ...:     mask[~np.in1d(np.arange(mask.shape[0]),A)] = 0
     ...:     R,C = np.where(mask)
     ...:     out = np.split(C,np.flatnonzero(R[1:]>R[:-1])+1)
     ...:     return out
     ...: 

In [136]: A = np.unique(np.random.randint(0,10000,(400)))
     ...: B = np.random.randint(0,10000,(400,300))
     ...: 

In [137]: %timeit index_1din2d_initbased(A,B)
10 loops, best of 3: 57.5 ms per loop

In [138]: %timeit index_1din2d_initbased_v2(A,B)
10 loops, best of 3: 25.9 ms per loop

网友

2楼 · 编辑于 2024-09-28 21:38:51

组合了numpy和list-comprehension的选项：

import numpy as np
[np.where((B == x).sum(axis = 1))[0] for x in A]
# [array([0, 1, 3]), array([1, 2, 3]), array([0, 1, 2]), array([0, 2, 3])]

相关问题更多 >

编程相关推荐

热门问题

热门文章

numpy返回一个3d数组内的索引

相关问题 更多 >

编程相关推荐

热门问题

热门文章

相关问题更多 >