My Account Log in

1 option

How to collect a corpus of websites with a web crawler / James A. Hodges.

SAGE Research Methods: Doing Research Online Available online

View online
Format:
Book
Author/Creator:
Hodges, James A., active 2022, author.
Series:
SAGE research methods: doing research online
Language:
English
Subjects (All):
Internet searching--Computer programs.
Internet searching.
Web archiving.
Physical Description:
1 online resource.
Place of Publication:
London : SAGE Publications, Ltd., 2022.
Summary:
Conducting research on digital cultures often requires some form of reference to online sources-but online sources are constantly changing, being updated, or deleted on a minute-by-minute basis. This guide will introduce the use of web crawlers as one potential method for gathering a stable, trustworthy collection of online sources. A corpus of sources generated via a web crawler can function as a detailed snapshot of the way an online resource existed at a particular point in time. The guide begins with an introduction to the theory behind web crawling, before moving into discussions of ethical concerns and commonly used tools. After addressing each of these foundational areas, the guide concludes with a step-by-step demonstration of web crawling with the popular command-line based open-source web downloading tool known as Wget.
Notes:
Description based on publisher supplied metadata and other sources.
ISBN:
1-5296-0932-1
9781529609325
OCLC:
1335036630

The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.

Find

Home Release notes

My Account

Shelf Request an item Bookmarks Fines and fees Settings

Guides

Using the Find catalog Using Articles+ Using your account