
Book Details
| Print ISBN | 9781394234493 |
| eText ISBN | 9781394234509 |
| Publisher | Wiley-Blackwell |
| Publishing Year | 2024 |
| Edition | 2nd Edition |
| Language | English |
| Pages | 368 |
The Data Science Handbook, 2nd Edition provides a broad technical grounding in core analytical methods for aspiring data scientists and professionals in non-data fields seeking to use analytics. Published by Wiley-Blackwell, this handbook connects foundational statistical concepts directly with practical software engineering practices across 368 pages of structured technical material.
The text synthesizes essential data operations into organized subject areas, starting with data munging, regular expressions, string manipulation, and metric visualizations. It then progresses into machine learning workflows, covering classification algorithms, feature extraction ideas, regression modeling, and unsupervised techniques including clustering and dimensionality reduction. Additional thematic sections detail traditional natural language processing, time series analysis, big data storage, database systems, probability, and inferential statistics.
All practical code examples across the volume rely on Python. The handbook includes focused sections on technical communication, programming language concepts, performance, computer memory management, data structures, and maximum-likelihood estimation to support applied computational learning across quantitative environments.
Table of Contents
Chapter 1: Introduction
- • 1.1 What Data Science Is and Isn’t
- • 1.2 This Book’s Slogan: Simple Models Are Easier to Work With
- • 1.3 How Is This Book Organized?
- • 1.4 How to Use This Book?
- • 1.5 Why Is It All in Python, Anyway?
- • 1.6 Example Code and Datasets
- • 1.7 Parting Words
Chapter 2: The Data Science Road Map
- • 2.1 Frame the Problem
- • 2.2 Understand the Data: Basic Questions
- • 2.3 Understand the Data: Data Wrangling
- • 2.4 Understand the Data: Exploratory Analysis
- • 2.5 Extract Features
- • 2.6 Model
- • 2.7 Present Results
- • 2.8 Deploy Code
- • 2.9 Iterating
- • 2.10 Glossary
Chapter 3: Programming Languages
- • 3.1 Why Use a Programming Language? What Are the Other Options?
- • 3.2 A Survey of Programming Languages for Data Science
- • 3.3 Where to Write Code
- • 3.4 Python Overview and Example Scripts
- • 3.5 Python Data Types
- • 3.6 GOTCHA: Hashable and Unhashable Types
- • 3.7 Functions and Control Structures
- • 3.8 Other Parts of Python
- • 3.9 Python’s Technical Libraries
- • 3.10 Other Python Resources
- • 3.11 Further Reading
- • 3.12 Glossary
Chapter 3a: Interlude: My Personal Toolkit
Chapter 4: Data Munging: String Manipulation, Regular Expressions, and Data Cleaning
- • 4.1 The Worst Dataset in the World
- • 4.2 How to Identify Pathologies
- • 4.3 Problems with Data Content
- • 4.4 Formatting Issues
- • 4.5 Example Formatting Script
- • 4.6 Regular Expressions
- • 4.7 Life in the Trenches
- • 4.8 Glossary
Chapter 5: Visualizations and Simple Metrics
- • 5.1 A Note on Python’s Visualization Tools
- • 5.2 Example Code
- • 5.3 Pie Charts
- • 5.4 Bar Charts
- • 5.5 Histograms
- • 5.6 Means, Standard Deviations, Medians, and Quantiles
- • 5.7 Boxplots
- • 5.8 Scatterplots
- • 5.9 Scatterplots with Logarithmic Axes
- • 5.10 Scatter Matrices
- • 5.11 Heatmaps
- • 5.12 Correlations
- • 5.13 Anscombe’s Quartet and the Limits of Numbers
- • 5.14 Time Series
- • 5.15 Further Reading
- • 5.16 Glossary
Chapter 6: Overview: Machine Learning and Artificial Intelligence
- • 6.1 Historical Context
- • 6.2 The Central Paradigm: Learning a Function from Example
- • 6.3 Machine Learning Data: Vectors and Feature Extraction
- • 6.4 Supervised, Unsupervised, and In-Between
- • 6.5 Training Data, Testing Data, and the Great Boogeyman of Overfitting
- • 6.6 Reinforcement Learning
- • 6.7 ML Models as Building Blocks for AI Systems
- • 6.8 ML Engineering as a New Job Role
- • 6.9 Further Reading
- • 6.10 Glossary
Chapter 7: Interlude: Feature Extraction Ideas
- • 7.1 Standard Features
- • 7.2 Features that Involve Grouping
- • 7.3 Preview of More Sophisticated Features
- • 7.4 You Get What You Measure: Defining the Target Variable
Chapter 8: Machine-Learning Classification
- • 8.1 What Is a Classifier, and What Can You Do with It?
- • 8.2 A Few Practical Concerns
- • 8.3 Binary Versus Multiclass
- • 8.4 Example Script
- • 8.5 Specific Classifiers
- • 8.6 Evaluating Classifiers
- • 8.7 Selecting Classification Cutoffs
- • 8.8 Further Reading
- • 8.9 Glossary
Chapter 9: Technical Communication and Documentation
- • 9.1 Several Guiding Principles
- • 9.2 Slide Decks
- • 9.3 Written Reports
- • 9.4 Speaking: What Has Worked for Me
- • 9.5 Code Documentation
- • 9.6 Further Reading
- • 9.7 Glossary
Chapter 10: Unsupervised Learning: Clustering and Dimensionality Reduction
- • 10.1 The Curse of Dimensionality
- • 10.2 Example: Eigenfaces for Dimensionality Reduction
- • 10.3 Principal Component Analysis and Factor Analysis
- • 10.4 Skree Plots and Understanding Dimensionality
- • 10.5 Factor Analysis
- • 10.6 Limitations of PCA
- • 10.7 Clustering
- • 10.8 Further Reading
- • 10.9 Glossary
Chapter 11: Regression
- • 11.1 Example: Predicting Diabetes Progression
- • 11.2 Fitting a Line with Least Squares
- • 11.3 Alternatives to Least Squares
- • 11.4 Fitting Nonlinear Curves
- • 11.5 Goodness of Fit: R 2 and Correlation
- • 11.6 Correlation of Residuals
- • 11.7 Linear Regression
- • 11.8 LASSO Regression and Feature Selection
- • 11.9 Further Reading
- • 11.10 Glossary
Chapter 12: Data Encodings and File Formats
- • 12.1 Typical File Format Categories
- • 12.2 CSV Files
- • 12.3 JSON Files
- • 12.4 XML Files
- • 12.5 HTML Files
- • 12.6 Tar Files
- • 12.7 GZip Files
- • 12.8 Zip Files
- • 12.9 Image Files: Rasterized, Vectorized, and/or Compressed
- • 12.10 It’s All Bytes at the End of the Day
- • 12.11 Integers
- • 12.12 Floats
- • 12.13 Text Data
- • 12.14 Further Reading
- • 12.15 Glossary
Chapter 13: Big Data
- • 13.1 What Is Big Data?
- • 13.2 When to Use – And not Use – Big Data
- • 13.3 Hadoop: The File System and the Processor
- • 13.4 Example PySpark Script
- • 13.5 Spark Overview
- • 13.6 Spark Operations
- • 13.7 PySpark Data Frames
- • 13.8 Two Ways to Run PySpark
- • 13.9 Configuring Spark
- • 13.10 Under the Hood
- • 13.11 Spark Tips and Gotchas
- • 13.12 The MapReduce Paradigm
- • 13.13 Performance Considerations
- • 13.14 Further Reading
- • 13.15 Glossary
Chapter 14: Databases
- • 14.1 Relational Databases and MySQL®
- • 14.2 Key–Value Stores
- • 14.3 Wide-Column Stores
- • 14.4 Document Stores
- • 14.5 Further Reading
- • 14.6 Glossary
Chapter 15: Software Engineering Best Practices
- • 15.1 Coding Style
- • 15.2 Version Control and Git for Data Scientists
- • 15.3 Testing Code
- • 15.4 Test-Driven Development
- • 15.5 AGILE Methodology
- • 15.6 Further Reading
- • 15.7 Glossary
Chapter 16: Traditional Natural Language Processing
- • 16.1 Do I Even Need NLP?
- • 16.2 The Great Divide: Language Versus Statistics
- • 16.3 Example: Sentiment Analysis on Stock Market Articles
- • 16.4 Software and Datasets
- • 16.5 Tokenization
- • 16.6 Central Concept: Bag-of-Words
- • 16.7 Word Weighting: TF-IDF
- • 16.8 n-Grams
- • 16.9 Stop Words
- • 16.10 Lemmatization and Stemming
- • 16.11 Synonyms
- • 16.12 Part of Speech Tagging
- • 16.13 Common Problems
- • 16.14 Advanced Linguistic NLP: Syntax Trees, Knowledge, and Understanding
- • 16.15 Further Reading
- • 16.16 Glossary
Chapter 17: Time Series Analysis
- • 17.1 Example: Predicting Wikipedia Page Views
- • 17.2 A Typical Workflow
- • 17.3 Time Series Versus Time-Stamped Events
- • 17.4 Resampling and Interpolation
- • 17.5 Smoothing Signals
- • 17.6 Logarithms and Other Transformations
- • 17.7 Trends and Periodicity
- • 17.8 Windowing
- • 17.9 Brainstorming Simple Features
- • 17.10 Better Features: Time Series as Vectors
- • 17.11 Fourier Analysis: Sometimes a Magic Bullet
- • 17.12 Time Series in Context: The Whole Suite of Features
- • 17.13 Further Reading
- • 17.14 Glossary
Chapter 18: Probability
- • 18.1 Flipping Coins: Bernoulli Random Variables
- • 18.2 Throwing Darts: Uniform Random Variables
- • 18.3 The Uniform Distribution and Pseudorandom Numbers
- • 18.4 Nondiscrete, Noncontinuous Random Variables
- • 18.5 Notation, Expectations, and Standard Deviation
- • 18.6 Dependence, Marginal, and Conditional Probability
- • 18.7 Understanding the Tails
- • 18.8 Binomial Distribution
- • 18.9 Poisson Distribution
- • 18.10 Normal Distribution
- • 18.11 Multivariate Gaussian
- • 18.12 Exponential Distribution
- • 18.13 Log-Normal Distribution
- • 18.14 Entropy
- • 18.15 Further Reading
- • 18.16 Glossary
Chapter 19: Statistics
- • 19.1 Statistics in Perspective
- • 19.2 Bayesian Versus Frequentist: Practical Tradeoffs and Differing Philosophies
- • 19.3 Hypothesis Testing: Key Idea and Example
- • 19.4 Multiple Hypothesis Testing
- • 19.5 Parameter Estimation
- • 19.6 Hypothesis Testing: t-Test
- • 19.7 Confidence Intervals
- • 19.8 Bayesian Statistics
- • 19.9 Naive Bayesian Statistics
- • 19.10 Bayesian Networks
- • 19.11 Choosing Priors: Maximum Entropy or Domain Knowledge
- • 19.12 Further Reading
- • 19.13 Glossary
Chapter 20: Programming Language Concepts
- • 20.1 Programming Paradigms
- • 20.2 Compilation and Interpretation
- • 20.3 Type Systems
- • 20.4 Further Reading
- • 20.5 Glossary
Chapter 21: Performance and Computer Memory
- • 21.1 A Word of Caution
- • 21.2 Example Script
- • 21.3 Algorithm Performance and Big-O Notation
- • 21.4 Some Classic Problems: Sorting a List and Binary Search
- • 21.5 Amortized Performance and Average Performance
- • 21.6 Two Principles: Reducing Overhead and Managing Memory
- • 21.7 Performance Tip: Use Numerical Libraries When Applicable
- • 21.8 Performance Tip: Delete Large Structures You Don’t Need
- • 21.9 Performance Tip: Use Built-In Functions When Possible
- • 21.10 Performance Tip: Avoid Superfluous Function Calls
- • 21.11 Performance Tip: Avoid Creating Large New Objects
- • 21.12 Further Reading
- • 21.13 Glossary
Chapter 22: Computer Memory and Data Structures
- • 22.1 Virtual Memory, the Stack, and the Heap
- • 22.2 Example C Program
- • 22.3 Data Types and Arrays in Memory
- • 22.4 Structs
- • 22.5 Pointers, the Stack, and the Heap
- • 22.6 Key Data Structures
- • 22.7 Further Reading
- • 22.8 Glossary
Chapter 23: Maximum-Likelihood Estimation and Optimization
- • 23.1 Maximum-Likelihood Estimation
- • 23.2 A Simple Example: Fitting a Line
- • 23.3 Another Exam
Customer Reviews
0.0
0 reviews
No reviews yet. Be the first to review this book!
Write a Review
Reviewed by GradeFocus Editorial Team
▶Research Sources (14)
- The Data Science Handbook, eBook by Field Cady | 9781394234509
- The Data Science Handbook eBook by Field Cady - EPUB | Rakuten ...
- The Data Science Handbook - Porrúa
- The data science handbook
- The data science handbook / Field Cady - Penn State University ...
- https://lernerbooks.com/products/search_results?se...
- Sourcebooks, LLC.
- Please verify you are human - Captcha
- Python Data Science Handbook
- Applied Data Sciences - Library Guides - Penn State
- Practical Data Science with Python or ...
- The Data Science Handbook (Hardcover) by Field Cady
- Python Data Science Handbook - 2nd Edition by Jake ...
- Python Data Science Handbook, 2nd Edition (Second ...




