获取复杂元素的属性使用lxm

<brandName type="http://example.com/codes/bmw#" abbrev="BMW" value="BMW" />BMW</brandName> <maxspeed> <value>250</value> <unit type="http://example.com/codes/units#" value="miles per hour" abbrev="mph" /> </maxspeed>

<Car xmlns="http://example.com/vocab/xml/cars#"> <dateStarted>2011-02-05</dateStarted> <dateSold>2011-02-13</dateSold> <name type="http://example.com/codes/bmw#" abbrev="X6" value="BMW X6" >BMW X6</name> <brandName type="http://example.com/codes/bmw#" abbrev="BMW" value="BMW" />BMW</brandName> <maxspeed> <value>250</value> <unit type="http://example.com/codes/units#" value="miles per hour" abbrev="mph" /> </maxspeed> <route type="http://example.com/codes/routes#" abbrev="HW" value="Highway" >Highway</route> <power> <value>180</value> <unit type="http://example.com/codes/units#" value="powerhorse" abbrev="ph" /> </power> <frequency type="http://example.com/codes/frequency#" value="daily" >Daily</frequency> </Car>

2条回答

网友

1楼 · 编辑于 2024-10-01 00:32:27

import lxml.etree as ET
content='''
<Car xmlns="http://example.com/vocab/xml/cars#">
 <brandName type="http://example.com/codes/bmw#" abbrev="BMW" value="BMW" >BMW</brandName>
   <maxspeed>
     <value>250</value>
     <unit type="http://example.com/codes/units#" value="miles per hour" abbrev="mph" />
   </maxspeed>
 </Car>
'''

doc=ET.fromstring(content)
NS = 'http://example.com/vocab/xml/cars#'
# print(ET.tostring(doc,pretty_print=True))
for x in doc.xpath('//ns:maxspeed/ns:unit/@abbrev',namespaces={'ns': NS}):
    print(x)

收益率

^{pr2}$

网友

2楼 · 编辑于 2024-10-01 00:32:27

lxml元素上的.find方法将只搜索该元素的直接子子级。例如，在这个xml中：

<root>
    <brandName type="http://example.com/codes/bmw#" abbrev="BMW" value="BMW">BMW</brandName>
    <maxspeed>
        <value>250</value>
        <unit type="http://example.com/codes/units#" value="miles per hour" abbrev="mph" />
    </maxspeed>
</root>

您可以使用根元素.find方法来定位brandname元素或maxspeed元素，但是搜索不会在这些内部元素中遍历。在

例如，你可以这样做：

^{pr2}$

从这个返回的元素可以访问属性。在

如果要搜索XML文档中的所有元素，可以使用.iter（）方法。对于前面的例子，您可以说：

for element in root.iter(tag='unit'):
    print element #This would print all the unit elements in the document.

编辑：下面是一个使用您提供的xml的全功能小示例：

import lxml.etree
from StringIO import StringIO

def ns_join(element, tag, namespace=None):
    '''Joins the namespace and tag together, and
    returns the fully qualified name.
    @param element - The lxml.etree._Element you're searching
    @param tag - The tag you're joining
    @param namespace - (optional) The Namespace shortname default is None'''

    return '{%s}%s' % (element.nsmap[namespace], tag)

def parse_car(element):
    '''Parse a car element, This will return a dictionary containing
    brand_name, maxspeed_value, and maxspeed_unit'''

    maxspeed = element.find(ns_join(element,'maxspeed'))
    return { 
        'brand_name' : element.findtext(ns_join(element,'brandName')), 
        'maxspeed_value' : maxspeed.findtext(ns_join(maxspeed,'value')), 
        'maxspeed_unit' : maxspeed.find(ns_join(maxspeed, 'unit')).attrib['abbrev']
        }

#Create the StringIO object to feed to the parser.
XML = StringIO('''
<Reports>
    <Car xmlns="http://example.com/vocab/xml/cars#">
        <dateStarted>2011-02-05</dateStarted>
        <dateSold>2011-02-13</dateSold>
        <name type="http://example.com/codes/bmw#" abbrev="X6" value="BMW X6" >BMW X6</name>
        <brandName type="http://example.com/codes/bmw#" abbrev="BMW" value="BMW" >BMW</brandName>
        <maxspeed>
            <value>250</value>
            <unit type="http://example.com/codes/units#" value="miles per hour" abbrev="mph" />
        </maxspeed>
        <route type="http://example.com/codes/routes#" abbrev="HW" value="Highway" >Highway</route>
        <power>
            <value>180</value>
            <unit type="http://example.com/codes/units#" value="powerhorse" abbrev="ph" />
        </power>
        <frequency type="http://example.com/codes/frequency#" value="daily" >Daily</frequency>  
    </Car>
</Reports>
''')

#Get the root element object of the xml
car_root_element = lxml.etree.parse(XML).getroot()

# For each 'Car' tag in the root element,
# we want to parse it and save the list as cars
cars = [ parse_car(element) 
    for element in car_root_element.iter() if element.tag.endswith('Car')]

print cars

希望有帮助。在

相关问题更多 >

编程相关推荐

热门问题

热门文章