<p>如果要使用<code>requests + Beautfulsoup</code>,请尝试以下操作(通过传递参数<code>page</code>):</p>
<pre><code>import re, requests, threading, os
from bs4 import BeautifulSoup
def download_image(url):
with open(os.path.basename(url), "wb") as f:
f.write(requests.get(url).content)
print(url, "download successfully")
original_url = "https://www.flickr.com/search/?text=sea&view_all=1&page={}"
pages = range(1, 5000) # not sure how many pages here
for page in pages:
concat_url = original_url.format(page)
print("Now it is page", page)
soup = BeautifulSoup(requests.get(concat_url).content, "lxml")
soup_list = soup.select(".photo-list-photo-view")
for element in soup_list:
img_url = 'https:'+re.search(r'url\((.*)\)', element.get("style")).group(1)
# the url like: https://live.staticflickr.com/xxx/xxxxx_m.jpg
# if you want to get a clearer(and larger) picture, remove the "_m" in the end of the url.
# For prevent IO block,I create a thread to download it.pass the url of the image as argument.
threading.Thread(target=download_image, args=(img_url,)).start()
</code></pre>
<hr/>
<p>如果使用selenium,可能会更简单,示例代码如下:</p>
<pre class="lang-py prettyprint-override"><code>from selenium import webdriver
import re, requests, threading, os
# download_image
def download_image(url):
with open(os.path.basename(url), "wb") as f:
f.write(requests.get(url).content)
driver = webdriver.Chrome()
original_url = "https://www.flickr.com/search/?text=sea&view_all=1&page={}"
pages = range(1, 5000) # not sure how many pages here
for page in pages:
concat_url = original_url.format(page)
print("Now it is page", page)
driver.get(concat_url)
for element in driver.find_elements_by_css_selector(".photo-list-photo-view"):
img_url = 'https:'+re.search(r'url\(\"(.*)\"\)', element.get_attribute("style")).group(1)
# the url like: https://live.staticflickr.com/xxx/xxxxx_m.jpg
# if you want to get a clearer(and larger) picture, remove the "_m" in the end of the url.
# For prevent IO block,I create a thread to download it.pass the url of the image as argument.
threading.Thread(target=download_image, args=(img_url, )).start()
</code></pre>
<p>并在我的电脑上成功下载</p>
<p><a href="https://i.stack.imgur.com/YWUMR.png" rel="nofollow noreferrer"><img src="https://i.stack.imgur.com/YWUMR.png" alt="enter image description here"/></a></p>