新的pandas数据帧来自另一个基于唯一多列索引的数据帧

import sqlite3 as db import pandas as pd conn = db.connect('C:/data.db') query = """SELECT TimeStamp, UnderlyingSymbol, Expiry, Strike, CP, BisectIV, OTMperc FROM ActiveOptions WHERE TimeStamp = '2015-11-09 16:00:00' AND UnderlyingSymbol = 'INTC' AND Expiry < '2015-11-27 16:00:00' AND OTMperc < .02 AND OTMperc > -.02 ORDER BY UnderlyingSymbol, Expiry, ABS(OTMperc)""" df = pd.read_sql_query(sql=query, con=conn,index_col=['TimeStamp', 'UnderlyingSymbol', 'Expiry'], parse_dates=['TimeStamp', 'Expiry'])

In[10]: new_df = df.index.drop_duplicates() In[11]: new_df Out[11]: MultiIndex(levels=[[2015-11-09 16:00:00], [u'INTC'], [2015-11-13 16:00:00, 2015-11-20 16:00:00]], labels=[[0, 0], [0, 0], [0, 1]], names=[u'TimeStamp', u'UnderlyingSymbol', u'Expiry']) In[12]: type(new_df) Out[12]: pandas.core.index.MultiIndex

1条回答

网友

1楼 · 发布于 2024-09-28 23:15:30

问题是您将new_df设置为索引列表，并删除了重复项：

new_df = df.index.drop_duplicates()

您需要的是只选择没有重复索引的行。您可以使用^{}函数筛选旧数据帧：

^{pr2}$

一个小例子，基于this：

#create data sample with multi index
arrays = [['bar', 'bar', 'baz', 'baz', 'foo', 'foo', 'qux', 'qux'],
          ['one', 'one', 'one', 'two', 'one', 'two', 'one', 'one']]
#(the first and last are duplicates)
tuples = list(zip(*arrays))
index = pd.MultiIndex.from_tuples(tuples, names=['first', 'second'])
s = pd.Series(np.random.randn(8), index=index)

原始数据：

>>> s
first  second
bar    one      -0.932521
       one       1.969771
baz    one       1.574908
       two       0.125159
foo    one      -0.075174
       two       0.777039
qux    one      -0.992862
       one      -1.099260
dtype: float64

并过滤重复项：

>>> s[~s.index.duplicated()]
first  second
bar    one      -0.932521
baz    one       1.574908
       two       0.125159
foo    one      -0.075174
       two       0.777039
qux    one      -0.992862
dtype: float64

相关问题更多 >

编程相关推荐

热门问题

热门文章