1 option
Mapping texts : computational text analysis for the social sciences / Dustin S. Stoltz and Marshall A. Taylor.
- Format:
- Book
- Author/Creator:
- Stoltz, Dustin S., author.
- Taylor, Marshall A., author.
- Series:
- Computational social science.
- Oxford scholarship online.
- Computational social science
- Oxford scholarship online
- Language:
- English
- Subjects (All):
- Text data mining.
- Physical Description:
- 1 online resource (326 pages)
- Place of Publication:
- New York, NY : Oxford University Press, 2024.
- Summary:
- 'Mapping Texts' provides an introduction to computational text analysis that simultaneously blends conceptual treatments with practical, hands-on examples that walk the reader through how to conduct text analysis projects with real data. The book shows how to conduct text analysis in the R statistical computing environment - a popular programming language in data science.
- Contents:
- Cover
- Advance Praise for Mapping Texts
- Mapping Texts: Computational Text Analysis for the Social Sciences
- Copyright
- Dediaction
- Contents
- Preface
- What You Will Learn
- What We Left Out
- Acknowledgments
- Part I: Bounding Texts
- 1: Text in Context
- What Is Language?
- What Is Text?
- 2: Corpus Building
- Texts Are Not People
- Balance, Range, and Representativeness
- Text Metadata
- Authors and Audiences
- Time and Location
- Domains and Media
- Text Data
- Languages and Dialects
- Genres and Topics
- Registers and Styles
- Redrawing Boundaries
- Part II: Prerequisites
- 3: Computing Basics
- Brass Tacks
- Coding Environments
- Data Objects, Types, and Structures
- Dialects of R
- Control Processes: Functions, Loops, and Apply
- Installing and Loading Packages
- Using Python in R
- Data Visualization
- Where to from Here
- 4: Math Basics
- The Fundamentals
- Comparing Vectors
- Dot Product
- Euclidean Distance and Cosine Similarity
- Correlation
- Regression
- Comparing Distributions
- Central Tendency
- Dispersion
- Types of Distributions
- Our Dear Friend, the Matrix
- Matrix Projection
- Vector Spaces and Singular Value Decomposition
- Graphs and Matrix Projection
- A Little Math Goes a Long Way
- Part III: Foundations
- 5: Acquiring Text
- Public Text Datasets
- Optical Character Recognition
- Automated Audio Transcription
- Application Programming Interfaces (APIs)
- Automated Web Scraping
- Legal and Ethical Side of Scraping
- Terms of Service
- Intellectual Property
- Individual and Organizational Privacy
- 6: From Text to Numbers
- Units of Analysis
- Tokenizing
- Chunking
- Document Features
- Sparsity
- Dedicated DTM Functions
- Token Distributions
- Zipf's Law and Herdan-heaps' Law
- Weighting and Norming
- Relative Term Frequency.
- Term Frequency/inverse Document Frequency
- Summary of Weightings
- Term Features
- Dimension Reduction
- Part IV: Below the Document
- 7: Wrangling Words
- Character Encoding
- Markup Characters
- Removing and Replacing Characters
- "misspelled" Words
- Removing Words and Stoplists
- Replacing Words
- Stemming
- Lemmatizing
- Lemmatizing in French
- Wrangling Workflow
- 8: Tagging Words
- Dictionary Tagging
- Named-Entity Recognition
- How Named-entity Recognition Works
- Named Entities in R
- Part-of-Speech and Dependency Parsing
- How Pos Tagging Works
- Part-of-speech Tagging in R
- Dependency Parsing
- Part-of-speech Tagging for French
- Part V: The Document and Beyond
- 9: Core Deductive
- Discrete Indicators
- Weighted Indicators
- Frequency-weighted
- Term-weghted Dictionaries
- Selecting and Building Dictionaries
- Pre-built Dictionaries
- Building Dictionaries With Supervised Learning
- 10: Core Inductive
- Document Similarity
- One-mode Projections
- Euclidean Distances and Cosine Similarities
- Document Clustering
- Hierarchical Clustering
- K-means Clustering
- Topic Modeling
- Topic Modeling With Lsa
- Topic Modeling With Lda
- 11: Extended Inductive
- Inference and Topic Models
- Predicting Topics With Covariates
- Topic Prevalence, Conditional on Covariates
- Topic Content, Conditional on Covariates
- Word Embeddings: The First Generation
- Word Embedding Basics
- Weighting the Tcm
- Dimension Reduction With Singular Value Decomposition
- Word Embeddings: The Next Generation
- The Global Approach: Glove
- The Neural Network Approach: Cbow, Sngs, and Fasttext
- Contextualized Embeddings: Elmo and Bert
- Inductive Analysis With Word Embeddings
- Semantic Change
- Semantic Directions and Semantic Centroids
- Word Mover's Distance
- 12: Extended Deductive.
- Supervision and Validation
- Representing Objects as Features
- Splitting Corpora
- Classic Training With Supervision
- Logistic Regression
- Naive Bayes
- Training With Neural Networks
- Deductive Analysis With Pretrained Models
- Neural Networks With Pretrained Embeddings
- Concept Mover's Distance With Pretrained Embeddings
- Retrofitting Pretrained Embeddings for Deductive Analysis
- Inference With Text Networks
- Building Text Networks
- Centrality
- Degree Centrality
- Betweenness Centrality
- Network Backbones
- Univariate Network Inference
- Multivariate Network Inference
- 13: Project Workflow and Iteration
- The Paradox of the Complete Map
- Containerize Our Projects
- Memoing and Datasheets
- Repeating, Replicating, and Simulating the Null
- Knowledge Takes a Village
- Appendix
- References
- Index.
- Notes:
- Also issued in print: 2024.
- Includes bibliographical references and index.
- Description based on online resource and publisher information; title from PDF title page (viewed on November 16, 2023).
- Other Format:
- Print version: Stoltz, Dustin S. Mapping Texts
- ISBN:
- 0-19-775691-3
- 0-19-775689-1
- 0-19-775690-5
- OCLC:
- 1409615563
The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.