我正试图在我的公司刮一个网页,并将结果写入一个CSV文件。你知道吗
我可以通过以下代码获得所需的数据:
page = requests.get('https://wiki.us.cworld.company.com/display/6TO/AWS+Accounts', auth=('tdunphy', 'secret!'))
soup = BeautifulSoup(page.text, 'html.parser')
html = list(soup.children)[1]
all_rows = soup.find_all('tr')
row_count = 0
for row in all_rows:
row_count += 1
if row_count == 1:
continue
print(row.get_text())
但最终的数据是一起运行的,几乎无法破译:
company-govcloud-ab-mc-stage-adminkpmg-us-aws-adv-ab-mc-govcloud-admin-stageCommercial AccountAdvisory12345678901NoIslandhttps://company-govcloud-ab-mc-stage-admin.signin.aws.amazon.com/consoleKarel Somebody23452126676371Console, Access Key
company-govcloud-ab-mc-stagekpmg-us-aws-adv-ab-mc-govcloud-stageGov AccountAdvisory12324546562NoIslandhttps://company-govcloud-ab-mc-stage.signin.amazonaws-us-gov.com/consoleKarel Somebody123213123131Console, Access Key
company-cob(Decommissioned 03/28/2019)company-COB COB, Client OnboardingAdvisory21234546789812NoIslandhttps://company-cob.signin.aws.amazon.com/console/Laurence LorcaPending DecommissionConsole, Access Key
我希望生成的CSV具有以下标题:
['Company Account Name', 'AWS Account Name', 'Description', 'LOB', 'AWS Account Number', 'CIDR Block', 'Connected to Montvale', 'Peninsula or Island', 'URL', 'Owner', 'Engagement Code', 'CloudOps Access Type']
在原始网页上,数据位于HTML表格中,结果清晰可见:
company-govcloud-ab-mc-stage-admin company-us-aws-adv-ab-mc-govcloud-admin-stage Commercial Account Advisory 12345667890101 No Island https://company-govcloud-ab-mc-stage-admin.signin.aws.amazon.com/console Karel Somebody 123456789101 Console, Access Key
下面是我提取的数据中的一些HTML示例:
<tr><td class="confluenceTd">company-master</td><td class="confluenceTd">us-ktawsmasacct</td><td class="confluenceTd">Master Account</td><td class="confluenceTd">BPG</td><td class="confluenceTd"><span style="text-decoration: none;">123456789101</span></td><td colspan="1" class="confluenceTd"><br/></td><td class="confluenceTd">No</td><td class="confluenceTd">N/A - no cloud resources</td><td class="confluenceTd"><a href="https://us-ktech-aws-master-acct.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://us-ktech-aws-master-acct.signin.aws.amazon.com/console</a></td><td class="confluenceTd">Rahul Arya</td><td class="confluenceTd"> 245612345678</td><td class="confluenceTd">Console, Access Key</td></tr><tr><td class="confluenceTd">company-transit-hub1</td><td class="confluenceTd">us-ktawsth1acct</td><td class="confluenceTd">Transit Hub</td><td class="confluenceTd">BPG</td><td class="confluenceTd"><span style="text-decoration: none;">303779310401</span></td><td colspan="1" class="confluenceTd"><span style="color: rgb(0,0,0);">10.47.0.0/24</span></td><td class="confluenceTd">No</td><td class="confluenceTd">Peninsula</td><td class="confluenceTd"><a href="https://company-transit-hub1.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://company-transit-hub1.signin.aws.amazon.com/console</a></td><td class="confluenceTd">Rahul Arya</td><td class="confluenceTd"> 245612345678</td><td class="confluenceTd">Console, Access Key</td></tr>
<tr><td colspan="1" class="confluenceTd">company-transit-hub3 (lab)</td><td colspan="1" class="confluenceTd"><span style="color: rgb(68,68,68);text-decoration: none;">us-dbawsth3acct</span></td><td colspan="1" class="confluenceTd">Transit Hub</td><td colspan="1" class="confluenceTd">BPG</td><td colspan="1" class="confluenceTd"><span style="color: rgb(68,68,68);text-decoration: none;">1098765432101</span> </td><td colspan="1" class="confluenceTd"><span style="color: rgb(0,0,0);">10.0.0.0/24</span></td><td colspan="1" class="confluenceTd">No</td><td colspan="1" class="confluenceTd">Island</td><td colspan="1" class="confluenceTd"> <a href="https://company-transithub3-lab.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://company-transithub3-lab.signin.aws.amazon.com/console</a></td><td colspan="1" class="confluenceTd">Rahul Arya </td><td colspan="1" class="confluenceTd"> </td><td colspan="1" class="confluenceTd">Console, Access Key</td></tr>
<tr><td class="confluenceTd">company-security</td><td class="confluenceTd"><span style="color: rgb(68,68,68);text-decoration: none;">us-ktawssecacct</span></td><td class="confluenceTd">Security</td><td class="confluenceTd">BPG</td><td class="confluenceTd">254312345691</td><td colspan="1" class="confluenceTd"><br/></td><td class="confluenceTd">No</td><td class="confluenceTd"><span>connected through hub1</span></td><td class="confluenceTd"><a href="https://us-ktawssecacct.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://us-ktawssecacct.signin.aws.amazon.com/console</a></td><td class="confluenceTd">Rahul Arya</td><td class="confluenceTd"> 245612345678</td><td class="confluenceTd">Console, Access Key</td></tr><tr><td class="confluenceTd">company-shared-services</td><td class="confluenceTd">us-ktawsssacct</td><td class="confluenceTd">Shared Services</td><td class="confluenceTd">BPG</td><td class="confluenceTd">300944922012</td><td colspan="1" class="confluenceTd"><br/></td><td class="confluenceTd">No</td><td class="confluenceTd"><span>connected through hub1</span></td><td class="confluenceTd"><a href="https://company-shared-services.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://company-shared-services.signin.aws.amazon.com/console</a></td><td class="confluenceTd">Rahul Arya</td><td class="confluenceTd">245612345678</td><td class="confluenceTd">Console, Access Key</td></tr><tr>
<tr><td class="confluenceTd">company-logging</td><td class="confluenceTd">us-ktawslogmonacct</td><td class="confluenceTd">Logging</td><td class="confluenceTd">BPG</td><td class="confluenceTd">542348765123</td><td colspan="1" class="confluenceTd"><br/></td><td class="confluenceTd">No</td><td class="confluenceTd"><span>connected through hub1</span></td><td class="confluenceTd"><a href="https://company-logging.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://company-logging.signin.aws.amazon.com/console</a></td><td class="confluenceTd">Rahul Arya</td><td class="confluenceTd">800000039768</td><td class="confluenceTd">Console, Access Key</td></tr><tr><td class="confluenceTd">company-spoke-acct1</td><td class="confluenceTd">us-ktawsspk1acct</td><td class="confluenceTd">Spoke Account</td><td class="confluenceTd">BPG</td><td class="confluenceTd"><span style="text-decoration: none;">103440952267</span></td><td colspan="1" class="confluenceTd"><span style="color: rgb(0,0,0);text-decoration: none;">10.47.8.0/24</span></td><td class="confluenceTd">No</td><td class="confluenceTd"><span>connected through hub1</span></td><td class="confluenceTd"><a href="https://block-chain.signin.aws.amazon.com/console" class="external-link" rel="nofollow">https://block-chain.signin.aws.amazon.com/console</a></td><td class="confluenceTd">Rahul Arya</td><td class="confluenceTd"><p>123456757897</p></td><td class="confluenceTd">Console, Access Key</td></tr>
问题是,当我从页面中刮取数据时,数据一起运行,我需要分离数据并插入逗号。你知道吗
如何在表数据的每个字段之间插入逗号,以便将其写入CSV文件?你知道吗
要写入CSV文件,请使用内置的
csv
模块:文件
out.csv
包含:LibreOffice Calc的屏幕截图:
相关问题 更多 >
编程相关推荐