无法获取相关链接并且排除其他链接

2024-04-27 08:03:08 发布

您现在位置:Python中文网/ 问答频道 /正文

我已经用python和selenium以及BeautifulSoup编写了一个脚本,从网页中获取指向属性详细信息的链接。由于内容是高度动态的,我使用了selenium来获取页面源代码。当我运行我的脚本时,我得到很多链接,包括那些必需的链接。你知道吗

如何仅从三个容器中的每个容器获取相关链接?

我的尝试:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def fetch_info(link):
    driver.get(link)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "#community-search-homes .propertyWrapper > a")))
    soup = BeautifulSoup(driver.page_source, "lxml")
    linklist = [item.get("href") for item in soup.select("#community-search-homes .propertyWrapper > a")]
    return linklist

if __name__ == '__main__':
    url = "https://www.khov.com/find-new-homes/arizona/buckeye"
    driver = webdriver.Chrome()
    wait = WebDriverWait(driver,10)
    for newlink in fetch_info(url):
        print(newlink)
    driver.quit()

结果是:

/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/aspire-at-sienna-hills
/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/affinity-at-verrado
/find-new-homes/arizona/buckeye/85396/four-seasons/k.-hovnanian's-four-seasons-at-victory-at-verrado
/find-new-homes/arizona/scottsdale/85255/k-hovnanian-homes/summit-at-silverstone
/find-new-homes/arizona/scottsdale/85257/k-hovnanian-homes/skye
/find-new-homes/arizona/phoenix/85020/k-hovnanian-homes/pointe-16
/find-new-homes/arizona/peoria/85383/k-hovnanian-homes/fusion-ii-at-the-meadows
/find-new-homes/arizona/scottsdale/85257/k-hovnanian-homes/aire
/find-new-homes/arizona/scottsdale/85255/k-hovnanian-homes/pinnacle-at-silverstone
/find-new-homes/arizona/peoria/85383/k-hovnanian-homes/montage-at-the-meadows
/find-new-homes/arizona/sun-city/85373/four-seasons/k.-hovnanian-s-four-seasons-at-ventana-lakes
/find-new-homes/arizona/peoria/85382/k-hovnanian-homes/park-paseo
/find-new-homes/arizona/laveen/85339/k-hovnanian-homes/affinity-at-montana-vista
/find-new-homes/arizona/laveen/85339/k-hovnanian-homes/aspire-at-montana-vista
/find-new-homes/arizona/scottsdale/85255/k-hovnanian-homes/pinnacle-ii-at-silverstone
/find-new-homes/arizona/scottsdale/85255/k-hovnanian-homes/summit-ii-at-silverstone

我想得到的结果是:

/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/aspire-at-sienna-hills
/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/affinity-at-verrado
/find-new-homes/arizona/buckeye/85396/four-seasons/k.-hovnanian's-four-seasons-at-victory-at-verrado

html元素块(the link I'm after is within the second line of the following elements):

<div class="propertyWrapper clear">
        <a href="/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/aspire-at-sienna-hills"><span class="link-outside"></span></a>
        <div class="propertyCarouselWrapper">
            <div class="responsiveImageCarousel enabled" style="touch-action: pan-y; user-select: none; -webkit-user-drag: none; -webkit-tap-highlight-color: rgba(0, 0, 0, 0);">
                <div class="prevBtn"></div>
                <div class="nextBtn"></div>
                <div class="images" data-detail-url="/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/aspire-at-sienna-hills">
                    <ul style="width: 960px; left: 0px;">
                        <li style="width: 320px;"><img alt="holiday exterior new homes sienna hills usp" src="https://khovcachecdn.azureedge.net/azure/sitefinitylibraries/images/default-source/images/az/aspire-at-sienna-hills/community-thumbnails/holiday-exterior-new-homes-sienna-hills-usp.jpg?sfvrsn=4&amp;build=1019&amp;encoder=wic&amp;useresizingpipeline=true&amp;w=450&amp;h=280&amp;mode=crop"></li>
                        <li style="width: 320px;"><img alt="carnival exterior new homes sienna hills usp" src="https://khovcachecdn.azureedge.net/azure/sitefinitylibraries/images/default-source/images/az/aspire-at-sienna-hills/community-thumbnails/carnival-exterior-new-homes-sienna-hills-usp.jpg?sfvrsn=4&amp;build=1019&amp;encoder=wic&amp;useresizingpipeline=true&amp;w=450&amp;h=280&amp;mode=crop"></li>
                    </ul>
                </div>
                <div class="pagination" style="width: 56px;"><ul><li class="active"></li><li></li></ul></div>
            </div>
        </div>
        <div class="propertyInfoWrapper">
            <div class="marker-details-container">
                <h3 class="marker-details">New Homes in Buckeye, Arizona</h3>
                <div class="spacer"></div>
                <h4 class="propertyListingHeader">Aspire at Sienna Hills</h4>
                <p class="marker-details">21007 West Almeria Road, Buckeye, AZ 85396</p>
                <p class="marker-details marker-status">Final Opportunities</p>
                <div class="spacer"></div>
                <p class="marker-details marker-price"><span class="bold">Priced from: </span>Mid $200s</p>
                <p class="marker-details"><span class="bold">Home type: </span>Single Family Homes</p>
                <p class="marker-details marker-amenities"><span class="bold">Amenities: </span>Pool, Hiking Trails, Park</p>
            </div>
            <div class="community-tag-container">
                <a href="/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/aspire-at-sienna-hills#quick-move-in-homes" onclick="KHOV.Analytics.trackEvent('Qmi_Icon_Qmi');">
                    <div class="community-tag">
                        <div class="ctaDesc quick-move-in-badge link-inside">Quick Move In Homes</div>
                        <div class="ctaIcon quick-move-in-badge-icon link-inside"></div>
                    </div>
                </a>
            </div>
            <a href="#request-info-form-modal" class="open-inline-modal-link" onclick="KHOV.Analytics.trackEvent('Orange_Ri_Request_Info');">
                <div class="button orange-color requestInfoButton link-inside" data-urlname="aspire-at-sienna-hills">Request Info</div>
            </a>

        </div>
    </div>

Tags: divnewlinkfindmarkeratclassamp
3条回答

您只需在链接中检查所需的关键字并打印这些关键字,然后忽略其他关键字:

if __name__ == '__main__':
    url = "https://www.khov.com/find-new-homes/arizona/buckeye"
    driver = webdriver.Chrome()
    wait = WebDriverWait(driver,10)
    for newlink in fetch_info(url):
        if url.split('/')[-1] in newlink:
            print(newlink)
    driver.quit()

输出:

/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/aspire-at-sienna-hills
/find-new-homes/arizona/buckeye/85396/k-hovnanian-homes/affinity-at-verrado
/find-new-homes/arizona/buckeye/85396/four-seasons/k.-hovnanian's-four-seasons-at-victory-at-verrado

会列出切片工作吗?你知道吗

def fetch_info(link):
    driver.get(link)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "#community-search-homes .propertyWrapper > a")))
    soup = BeautifulSoup(driver.page_source, "lxml")
    linklist = [item.get("href") for item in soup.select("#community-search-homes .propertyWrapper > a")][:3]
    return linklist

你需要包括特色id以及结果。您可以使用或组合。最新的bs4支持not。你知道吗

#propertyResultsContainer .propertyWrapper :not([onclick])[href*=find], #propertyFeaturedResultsContainer  .propertyWrapper :not([onclick])[href*=find]

这也可以缩短为

#propertyResultsContainer .propertyWrapper :not([onclick])[href*=find], #propertyFeaturedResultsContainer

但这种缩短可能不那么有力。你知道吗

相关问题 更多 >