Welcome to OStack Knowledge Sharing Community for programmer and developer-Open, Learning and Share
Welcome To Ask or Share your Answers For Others

Categories

0 votes
4.2k views
in Technique[技术] by (71.8m points)

web scraping - Python Scrapy - parse URL content for most recent updated date

I have a webcrawler/scraper that's written in Python using the scrapy framework. I've been trying to use the "last-modified" date to identify the most recent update for each page - but I also collect each HTML file for the pages that are scraped. Is there a more accurate method for collecting the date each page was most recently updated?


与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
Welcome To Ask or Share your Answers For Others

1 Answer

0 votes
by (71.8m points)
等待大神答复

与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
Welcome to OStack Knowledge Sharing Community for programmer and developer-Open, Learning and Share
Click Here to Ask a Question

...