如何将xml元素从一个非常大的xml文件解析为python?

2024-10-02 12:26:43 发布

您现在位置:Python中文网/ 问答频道 /正文

我目前正在开发一个程序,它有20个左右的脚本,可以从一个python文件调用,这个文件使用子进程库来调用这些脚本。每个脚本都有用户当前使用argparse输入的3个参数:ip地址、用户名和密码。这些脚本自动测试网络设备等

现在,我不想让用户在命令行中输入这些参数,而是想从一个XML文件中提取这些值,该文件包含我公司生成的大约5000行代码。我提取所需信息的最佳方式是什么,这样用户就不必手动输入参数了

我做了一些研究,不幸的是,我不能理解最好的方法来做这件事。以下是xml文件的示例摘录:

<sheet>
        <name>7_managementHosts</name>
        <data>
            <name>MgtHosts</name>
            <key>
                <name>Rack U-Location</name>
                <value>U30</value>
                <value>U29</value>
                <value>U28</value>
            </key>
            <key>
                <name>Default Component Name</name>
                <value>sms01</value>
                <value>sms02</value>
                <value>sms03</value>
            </key>
            <key>
                <name>DNS hostname (FQDN)</name>
                <value>sms01.de1000.local</value>
                <value>sms02.de1000.local</value>
                <value>sms03.de1000.local</value>
            </key>
            <key>
                <name>DNS suffix for management interface</name>
                <value>de1000.local</value>
                <value>de1000.local</value>
                <value>de1000.local</value>
            </key>
            <key>
                <name>Keyboard layout</name>
                <value>US Default</value>
                <value>US Default</value>
                <value>US Default</value>
            </key>
            <key>
                <name>root user password</name>
                <value>myPassword</value>
                <value>myPassword</value>
                <value>myPassword</value>
            </key>

这是一个很长的XML文件,但是树是这样的,我真的不知道最好的方法。谢谢你的帮助


Tags: 文件方法key用户name脚本default参数
2条回答

BeautifulSoup为例,让您从模块开始:

data = '''
<sheet>
        <name>7_managementHosts</name>
        <data>
            <name>MgtHosts</name>
            <key>
                <name>Rack U-Location</name>
                <value>U30</value>
                <value>U29</value>
                <value>U28</value>
            </key>
            <key>
                <name>Default Component Name</name>
                <value>sms01</value>
                <value>sms02</value>
                <value>sms03</value>
            </key>
            <key>
                <name>DNS hostname (FQDN)</name>
                <value>sms01.de1000.local</value>
                <value>sms02.de1000.local</value>
                <value>sms03.de1000.local</value>
            </key>
            <key>
                <name>DNS suffix for management interface</name>
                <value>de1000.local</value>
                <value>de1000.local</value>
                <value>de1000.local</value>
            </key>
            <key>
                <name>Keyboard layout</name>
                <value>US Default</value>
                <value>US Default</value>
                <value>US Default</value>
            </key>
            <key>
                <name>root user password</name>
                <value>myPassword</value>
                <value>myPassword</value>
                <value>myPassword</value>
            </key>
 '''

from bs4 import BeautifulSoup

data = BeautifulSoup(data, 'lxml')

parsed = [[v.text for v in key.select('name, value')] for key in data.select('key')]

# just for pretty printing, all the data are in `parsed` variable
from textwrap import shorten
for row_num, row in enumerate(zip(*parsed), 0):
    if row_num == 0:
        print(''.join('{: ^25}'.format(shorten(d, 25)) for d in ['Row Number'] + list(row)))
    else:
        print(''.join('{: ^25}'.format(shorten(d, 25)) for d in [str(row_num)] + list(row)))

印刷品:

   Row Number             Rack U-Location      Default Component Name     DNS hostname (FQDN)     DNS suffix for [...]        Keyboard layout        root user password    
        1                       U30                     sms01             sms01.de1000.local          de1000.local              US Default               myPassword        
        2                       U29                     sms02             sms02.de1000.local          de1000.local              US Default               myPassword        
        3                       U28                     sms03             sms03.de1000.local          de1000.local              US Default               myPassword        

使用pythonstandard XML lib(假设您想收集'key'元素下的数据)

import xml.etree.ElementTree as ET
import pprint

xml = '''<sheet>
        <name>7_managementHosts</name>
        <data>
            <name>MgtHosts</name>
            <key>
                <name>Rack U-Location</name>
                <value>U30</value>
                <value>U29</value>
                <value>U28</value>
            </key>
            <key>
                <name>Default Component Name</name>
                <value>sms01</value>
                <value>sms02</value>
                <value>sms03</value>
            </key>
            <key>
                <name>DNS hostname (FQDN)</name>
                <value>sms01.de1000.local</value>
                <value>sms02.de1000.local</value>
                <value>sms03.de1000.local</value>
            </key>
            <key>
                <name>DNS suffix for management interface</name>
                <value>de1000.local</value>
                <value>de1000.local</value>
                <value>de1000.local</value>
            </key>
            <key>
                <name>Keyboard layout</name>
                <value>US Default</value>
                <value>US Default</value>
                <value>US Default</value>
            </key>
            <key>
                <name>root user password</name>
                <value>myPassword</value>
                <value>myPassword</value>
                <value>myPassword</value>
            </key>
        </data>
    </sheet>'''

data = {}
root = ET.fromstring(xml)
keys = root.findall('.//data/key')
for key in keys:
    data[key.find('name').text] = [v.text for v in  key.findall('value')]
pprint.pprint(data)

输出

{'DNS hostname (FQDN)': ['sms01.de1000.local',
                         'sms02.de1000.local',
                         'sms03.de1000.local'],
 'DNS suffix for management interface': ['de1000.local',
                                         'de1000.local',
                                         'de1000.local'],
 'Default Component Name': ['sms01', 'sms02', 'sms03'],
 'Keyboard layout': ['US Default', 'US Default', 'US Default'],
 'Rack U-Location': ['U30', 'U29', 'U28'],
 'root user password': ['myPassword', 'myPassword', 'myPassword']}

相关问题 更多 >

    热门问题