我尝试使用beauthoulsoup解析以下HTML页面(我将解析大量页面)。在
我需要保存每个页面中的所有字段,但它们可以动态更改(在不同的页面上)。在
下面是一个页面-Page 1的示例 以及具有不同字段顺序的页-Page 2
我写了以下代码来解析这个页面。在
import requests
from bs4 import BeautifulSoup
PTiD = 7680560
url = "http://patft.uspto.gov/netacgi/nph-Parser?Sect1=PTO1&Sect2=HITOFF&d=PALL&p=1&u=/netahtml/PTO/srchnum.htm&r=1&f=G&l=50&s1=" + str(PTiD) + ".PN.&OS=PN/" + str(PTiD) + "&RS=PN/" + str(PTiD)
res = requests.get(url, prefetch = True)
raw_html = res.content
print "Parser Started.. "
bs_html = BeautifulSoup(raw_html, "lxml")
#Initialize all the Search Lists
fonts = bs_html.find_all('font')
para = bs_html.find_all('p')
bs_text = bs_html.find_all(text=True)
onlytext = [x for x in bs_text if x != '\n' and x != ' ']
#Initialize the Indexes
AppNumIndex = onlytext.index('Appl. No.:\n')
FiledIndex = onlytext.index('Filed:\n ')
InventorsIndex = onlytext.index('Inventors: ')
AssigneeIndex = onlytext.index('Assignee:')
ClaimsIndex = onlytext.index('Claims')
DescriptionIndex = onlytext.index(' Description')
CurrentUSClassIndex = onlytext.index('Current U.S. Class:')
CurrentIntClassIndex = onlytext.index('Current International Class: ')
PrimaryExaminerIndex = onlytext.index('Primary Examiner:')
AttorneyOrAgentIndex = onlytext.index('Attorney, Agent or Firm:')
RefByIndex = onlytext.index('[Referenced By]')
#~~Title~~
for a in fonts:
if a.has_key('size') and a['size'] == '+1':
d_title = a.string
print "title: " + d_title
#~~Abstract~~~
d_abstract = para[0].string
print "abstract: " + d_abstract
#~~Assignee Name~~
d_assigneeName = onlytext[AssigneeIndex +1]
print "as name: " + d_assigneeName
#~~Application number~~
d_appNum = onlytext[AppNumIndex + 1]
print "ap num: " + d_appNum
#~~Application date~~
d_appDate = onlytext[FiledIndex + 1]
print "ap date: " + d_appDate
#~~ Patent Number~~
d_PatNum = onlytext[0].split(':')[1].strip()
print "patnum: " + d_PatNum
#~~Issue Date~~
d_IssueDate = onlytext[10].strip('\n')
print "issue date: " + d_IssueDate
#~~Inventors Name~~
d_InventorsName = ''
for x in range(InventorsIndex+1, AssigneeIndex, 2):
d_InventorsName += onlytext[x]
print "inv name: " + d_InventorsName
#~~Inventors City~~
d_InventorsCity = ''
for x in range(InventorsIndex+2, AssigneeIndex, 2):
d_InventorsCity += onlytext[x].split(',')[0].strip().strip('(')
d_InventorsCity = d_InventorsCity.strip(',').strip().strip(')')
print "inv city: " + d_InventorsCity
#~~Inventors State~~
d_InventorsState = ''
for x in range(InventorsIndex+2, AssigneeIndex, 2):
d_InventorsState += onlytext[x].split(',')[1].strip(')').strip() + ','
d_InventorsState = d_InventorsState.strip(',').strip()
print "inv state: " + d_InventorsState
#~~ Asignee City ~~
d_AssigneeCity = onlytext[AssigneeIndex + 2].split(',')[1].strip().strip('\n').strip(')')
print "asign city: " + d_AssigneeCity
#~~ Assignee State~~
d_AssigneeState = onlytext[AssigneeIndex + 2].split(',')[0].strip('\n').strip().strip('(')
print "asign state: " + d_AssigneeState
#~~Current US Class~~
d_CurrentUSClass = ''
for x in range (CuurentUSClassIndex + 1, CurrentIntClassIndex):
d_CurrentUSClass += onlytext[x]
print "cur us class: " + d_CurrentUSClass
#~~ Current Int Class~~
d_CurrentIntlClass = onlytext[CurrentIntClassIndex +1]
print "cur intl class: " + d_CurrentIntlClass
#~~~Primary Examiner~~~
d_PrimaryExaminer = onlytext[PrimaryExaminerIndex +1]
print "prim ex: " + d_PrimaryExaminer
#~~d_AttorneyOrAgent~~
d_AttorneyOrAgent = onlytext[AttorneyOrAgentIndex +1]
print "agent: " + d_AttorneyOrAgent
#~~ Referenced by ~~
for x in range(RefByIndex + 2, RefByIndex + 400):
if (('Foreign' in onlytext[x]) or ('Primary' in onlytext[x])):
break
else:
d_ReferencedBy += onlytext[x]
print "ref by: " + d_ReferencedBy
#~~Claims~~
d_Claims = ''
for x in range(ClaimsIndex , DescriptionIndex):
d_Claims += onlytext[x]
print "claims: " + d_Claims
我将页面中的所有文本插入到一个列表中(使用beauthoulsoup的find_all(text=True))。然后,我试图找到字段名的索引,从该位置查看列表并将成员保存到字符串中,直到到达下一个字段索引为止。在
当我在几个不同的页面上尝试代码时,我注意到成员的结构正在发生变化,我在列表中找不到他们的索引。 例如,我搜索索引“123”,在某些页面上,它在列表中显示为“12”、“3”。在
你能想出其他方法来解析这个页面吗?在
谢谢。在
我认为最简单的解决方案是使用pyquery库 http://packages.python.org/pyquery/api.html
您可以使用jquery选择器选择页面的元素。在
如果您使用beauthoulsoup,并且有dom},那么您将拥有{}
<p>123</p>
和{但是,如果您有dom
<p>12<b>3</b></p>
,它与前面的语义相同,但是beautifulsoup将给您['12','3']
也许您可以准确地找到阻碍您完成
['123']
的标记,然后先忽略/消除该标记。在一些关于如何消除<;b>;标记的伪代码
对于图案,可以使用以下选项:
^{pr2}$相关问题 更多 >
编程相关推荐