使用BeautifulSoup抓取网页

import requests from bs4 import BeautifulSoup URL = 'https://www.senate.gov/general/contact_information/senators_cfm.cfm' page = requests.get(URL) soup = BeautifulSoup(page.content, 'html.parser') print(soup)

2条回答

网友

1楼 · 编辑于 2024-07-07 07:43:15

重复HTTP 503 Error while using python requests module

试试看：

import requests
from bs4 import BeautifulSoup

URL = 'https://www.senate.gov/general/contact_information/senators_cfm.cfm'

page = requests.post(URL, headers=headers)

soup = BeautifulSoup(page.content, 'html.parser')
print(soup)

网友

2楼 · 编辑于 2024-07-07 07:43:15

这对我有用

headers = {
        'user-agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.87 Safari/537.36',
    }
r = requests.get(URL,headers=headers)

在此处找到信息-https://towardsdatascience.com/5-strategies-to-write-unblock-able-web-scrapers-in-python-5e40c147bdaf

编程相关推荐

java缓存使用Spring Oauth2访问令牌
java我正在尝试在选项卡上插入一个编辑文本框
SwingJava：是否存在SwingGutilities。invokeNowOrLaterIfEDT（…）还是类似的？
关于Spring数据JDBC+Hikari+PostgresJSONB的java问题
java是否可以为tomcat数据源设置多个URL
java如何安排ServletContextListener在指定的日期和时间执行
javascript有没有办法使用本机应用程序的startActivityForResult从我开发到Android的PWA中获取结果？
java为什么建议将实例变量声明为私有？
java spring启动远程更新失败
安卓 java lang Null Pointerexception更新谷歌服务后

相关问题更多 >

编程相关推荐

热门问题

热门文章

使用BeautifulSoup抓取网页

相关问题 更多 >

编程相关推荐

热门问题

热门文章

相关问题更多 >