从pandas datafram将字符串数组（category）转换为int数组

trainedData=bigdata[bigdata['meta']<15] untrained=bigdata[bigdata['meta']>=15] #print trainedData #extract two columns from trainedData #convert to numpy array features=trainedData.ix[:,['ratio','area']].as_matrix(['ratio','area']) un_features=untrained.ix[:,['ratio','area']].as_matrix(['ratio','area']) print 'features' print features[:5] ##label is a string:single, touching,nuclei,dust print 'labels' labels=trainedData.ix[:,['type']].as_matrix(['type']) print labels[:5] #convert single to 0, touching to 1, nuclei to 2, dusts to 3 # tmp=categorical(labels,drop=True) targets=categorical(labels,drop=True).argmax(1) print targets

Traceback (most recent call last): File "/home/claire/Applications/ProjetPython/projet particule et objet/karyotyper/DAPI-Trainer02-MILK.py", line 83, in <module> tmp=categorical(labels,drop=True) File "/usr/local/lib/python2.6/dist-packages/scikits.statsmodels-0.3.0rc1-py2.6.egg/scikits/statsmodels/tools/tools.py", line 206, in categorical tmp_dummy = (tmp_arr[:,None]==data).astype(float) AttributeError: 'bool' object has no attribute 'astype'

3条回答

网友

1楼 · 编辑于 2024-06-26 10:51:58

我在回答熊猫0.10.1的问题。Factor.from_array似乎起到了作用。

>>> s = pandas.Series(['a', 'b', 'a', 'c', 'a', 'b', 'a'])
>>> s
0    a
1    b
2    a
3    c
4    a
5    b
6    a
>>> f = pandas.Factor.from_array(s)
>>> f
Categorical: 
array([a, b, a, c, a, b, a], dtype=object)
Levels (3): Index([a, b, c], dtype=object)
>>> f.labels
array([0, 1, 0, 2, 0, 1, 0])
>>> f.levels
Index([a, b, c], dtype=object)

网友

2楼 · 编辑于 2024-06-26 10:51:58

前面的答案已经过时了，所以这里有一个将字符串映射到数字的解决方案，它适用于0.18.1版的Pandas。

对于一个系列：

In [1]: import pandas as pd
In [2]: s = pd.Series(['single', 'touching', 'nuclei', 'dusts',
                       'touching', 'single', 'nuclei'])
In [3]: s_enc = pd.factorize(s)
In [4]: s_enc[0]
Out[4]: array([0, 1, 2, 3, 1, 0, 2])
In [5]: s_enc[1]
Out[5]: Index([u'single', u'touching', u'nuclei', u'dusts'], dtype='object')

对于数据帧：

In [1]: import pandas as pd
In [2]: df = pd.DataFrame({'labels': ['single', 'touching', 'nuclei', 
                       'dusts', 'touching', 'single', 'nuclei']})
In [3]: catenc = pd.factorize(df['labels'])
In [4]: catenc
Out[4]: (array([0, 1, 2, 3, 1, 0, 2]), 
        Index([u'single', u'touching', u'nuclei', u'dusts'],
        dtype='object'))
In [5]: df['labels_enc'] = catenc[0]
In [6]: df
Out[4]:
         labels  labels_enc
    0    single           0
    1  touching           1
    2    nuclei           2
    3     dusts           3
    4  touching           1
    5    single           0
    6    nuclei           2

网友

3楼 · 编辑于 2024-06-26 10:51:58

如果您有一个字符串或其他对象的向量，并且希望给它分类标签，那么可以使用Factor类（在pandas命名空间中可用）：

In [1]: s = Series(['single', 'touching', 'nuclei', 'dusts', 'touching', 'single', 'nuclei'])

In [2]: s
Out[2]: 
0    single
1    touching
2    nuclei
3    dusts
4    touching
5    single
6    nuclei
Name: None, Length: 7

In [4]: Factor(s)
Out[4]: 
Factor:
array([single, touching, nuclei, dusts, touching, single, nuclei], dtype=object)
Levels (4): [dusts nuclei single touching]

因子具有属性labels和levels：

In [7]: f = Factor(s)

In [8]: f.labels
Out[8]: array([2, 3, 1, 0, 3, 2, 1], dtype=int32)

In [9]: f.levels
Out[9]: Index([dusts, nuclei, single, touching], dtype=object)

这是针对一维向量的，所以不确定它是否可以立即应用到您的问题上，但请看一看。

顺便说一句，我建议你在statsmodels和/或scikit learn邮件列表上提出这些问题，因为我们大多数人都不经常使用。

相关问题更多 >

编程相关推荐

热门问题

热门文章