GradeFocus
BooksCategoriesAuthorsAboutContact
GradeFocus

Find textbooks and academic resources at competitive prices. Compare listings from VitalSource, Amazon, and more to save money on your course materials.

Browse

  • Books
  • Categories
  • Authors

Company

  • About
  • Contact
  • FAQ

Legal

  • Privacy
  • Terms
  • DMCA

© 2026 GradeFocus. All rights reserved.

PrivacyTermsSitemap
  1. Home
  2. /Data Science
The Data Science Handbook cover

The Data Science Handbook

by Field Cady

2nd Edition

Publisher: Wiley-Blackwell

(0 reviews)
Data Science

Compare Prices

VitalSourceLifetime Access$62.00AmazonKindle$62.00Best PriceeTextShelfPDF$38.00

Book Details

Print ISBN9781394234493
eText ISBN9781394234509
PublisherWiley-Blackwell
Publishing Year2024
Edition2nd Edition
LanguageEnglish
Pages368

The Data Science Handbook, 2nd Edition provides a broad technical grounding in core analytical methods for aspiring data scientists and professionals in non-data fields seeking to use analytics. Published by Wiley-Blackwell, this handbook connects foundational statistical concepts directly with practical software engineering practices across 368 pages of structured technical material.

The text synthesizes essential data operations into organized subject areas, starting with data munging, regular expressions, string manipulation, and metric visualizations. It then progresses into machine learning workflows, covering classification algorithms, feature extraction ideas, regression modeling, and unsupervised techniques including clustering and dimensionality reduction. Additional thematic sections detail traditional natural language processing, time series analysis, big data storage, database systems, probability, and inferential statistics.

All practical code examples across the volume rely on Python. The handbook includes focused sections on technical communication, programming language concepts, performance, computer memory management, data structures, and maximum-likelihood estimation to support applied computational learning across quantitative environments.

Table of Contents

  1. Chapter 1: Introduction

    • • 1.1 What Data Science Is and Isn’t
    • • 1.2 This Book’s Slogan: Simple Models Are Easier to Work With
    • • 1.3 How Is This Book Organized?
    • • 1.4 How to Use This Book?
    • • 1.5 Why Is It All in Python, Anyway?
    • • 1.6 Example Code and Datasets
    • • 1.7 Parting Words
  2. Chapter 2: The Data Science Road Map

    • • 2.1 Frame the Problem
    • • 2.2 Understand the Data: Basic Questions
    • • 2.3 Understand the Data: Data Wrangling
    • • 2.4 Understand the Data: Exploratory Analysis
    • • 2.5 Extract Features
    • • 2.6 Model
    • • 2.7 Present Results
    • • 2.8 Deploy Code
    • • 2.9 Iterating
    • • 2.10 Glossary
  3. Chapter 3: Programming Languages

    • • 3.1 Why Use a Programming Language? What Are the Other Options?
    • • 3.2 A Survey of Programming Languages for Data Science
    • • 3.3 Where to Write Code
    • • 3.4 Python Overview and Example Scripts
    • • 3.5 Python Data Types
    • • 3.6 GOTCHA: Hashable and Unhashable Types
    • • 3.7 Functions and Control Structures
    • • 3.8 Other Parts of Python
    • • 3.9 Python’s Technical Libraries
    • • 3.10 Other Python Resources
    • • 3.11 Further Reading
    • • 3.12 Glossary
  4. Chapter 3a: Interlude: My Personal Toolkit

  5. Chapter 4: Data Munging: String Manipulation, Regular Expressions, and Data Cleaning

    • • 4.1 The Worst Dataset in the World
    • • 4.2 How to Identify Pathologies
    • • 4.3 Problems with Data Content
    • • 4.4 Formatting Issues
    • • 4.5 Example Formatting Script
    • • 4.6 Regular Expressions
    • • 4.7 Life in the Trenches
    • • 4.8 Glossary
  6. Chapter 5: Visualizations and Simple Metrics

    • • 5.1 A Note on Python’s Visualization Tools
    • • 5.2 Example Code
    • • 5.3 Pie Charts
    • • 5.4 Bar Charts
    • • 5.5 Histograms
    • • 5.6 Means, Standard Deviations, Medians, and Quantiles
    • • 5.7 Boxplots
    • • 5.8 Scatterplots
    • • 5.9 Scatterplots with Logarithmic Axes
    • • 5.10 Scatter Matrices
    • • 5.11 Heatmaps
    • • 5.12 Correlations
    • • 5.13 Anscombe’s Quartet and the Limits of Numbers
    • • 5.14 Time Series
    • • 5.15 Further Reading
    • • 5.16 Glossary
  7. Chapter 6: Overview: Machine Learning and Artificial Intelligence

    • • 6.1 Historical Context
    • • 6.2 The Central Paradigm: Learning a Function from Example
    • • 6.3 Machine Learning Data: Vectors and Feature Extraction
    • • 6.4 Supervised, Unsupervised, and In-Between
    • • 6.5 Training Data, Testing Data, and the Great Boogeyman of Overfitting
    • • 6.6 Reinforcement Learning
    • • 6.7 ML Models as Building Blocks for AI Systems
    • • 6.8 ML Engineering as a New Job Role
    • • 6.9 Further Reading
    • • 6.10 Glossary
  8. Chapter 7: Interlude: Feature Extraction Ideas

    • • 7.1 Standard Features
    • • 7.2 Features that Involve Grouping
    • • 7.3 Preview of More Sophisticated Features
    • • 7.4 You Get What You Measure: Defining the Target Variable
  9. Chapter 8: Machine-Learning Classification

    • • 8.1 What Is a Classifier, and What Can You Do with It?
    • • 8.2 A Few Practical Concerns
    • • 8.3 Binary Versus Multiclass
    • • 8.4 Example Script
    • • 8.5 Specific Classifiers
    • • 8.6 Evaluating Classifiers
    • • 8.7 Selecting Classification Cutoffs
    • • 8.8 Further Reading
    • • 8.9 Glossary
  10. Chapter 9: Technical Communication and Documentation

    • • 9.1 Several Guiding Principles
    • • 9.2 Slide Decks
    • • 9.3 Written Reports
    • • 9.4 Speaking: What Has Worked for Me
    • • 9.5 Code Documentation
    • • 9.6 Further Reading
    • • 9.7 Glossary
  11. Chapter 10: Unsupervised Learning: Clustering and Dimensionality Reduction

    • • 10.1 The Curse of Dimensionality
    • • 10.2 Example: Eigenfaces for Dimensionality Reduction
    • • 10.3 Principal Component Analysis and Factor Analysis
    • • 10.4 Skree Plots and Understanding Dimensionality
    • • 10.5 Factor Analysis
    • • 10.6 Limitations of PCA
    • • 10.7 Clustering
    • • 10.8 Further Reading
    • • 10.9 Glossary
  12. Chapter 11: Regression

    • • 11.1 Example: Predicting Diabetes Progression
    • • 11.2 Fitting a Line with Least Squares
    • • 11.3 Alternatives to Least Squares
    • • 11.4 Fitting Nonlinear Curves
    • • 11.5 Goodness of Fit: R 2 and Correlation
    • • 11.6 Correlation of Residuals
    • • 11.7 Linear Regression
    • • 11.8 LASSO Regression and Feature Selection
    • • 11.9 Further Reading
    • • 11.10 Glossary
  13. Chapter 12: Data Encodings and File Formats

    • • 12.1 Typical File Format Categories
    • • 12.2 CSV Files
    • • 12.3 JSON Files
    • • 12.4 XML Files
    • • 12.5 HTML Files
    • • 12.6 Tar Files
    • • 12.7 GZip Files
    • • 12.8 Zip Files
    • • 12.9 Image Files: Rasterized, Vectorized, and/or Compressed
    • • 12.10 It’s All Bytes at the End of the Day
    • • 12.11 Integers
    • • 12.12 Floats
    • • 12.13 Text Data
    • • 12.14 Further Reading
    • • 12.15 Glossary
  14. Chapter 13: Big Data

    • • 13.1 What Is Big Data?
    • • 13.2 When to Use – And not Use – Big Data
    • • 13.3 Hadoop: The File System and the Processor
    • • 13.4 Example PySpark Script
    • • 13.5 Spark Overview
    • • 13.6 Spark Operations
    • • 13.7 PySpark Data Frames
    • • 13.8 Two Ways to Run PySpark
    • • 13.9 Configuring Spark
    • • 13.10 Under the Hood
    • • 13.11 Spark Tips and Gotchas
    • • 13.12 The MapReduce Paradigm
    • • 13.13 Performance Considerations
    • • 13.14 Further Reading
    • • 13.15 Glossary
  15. Chapter 14: Databases

    • • 14.1 Relational Databases and MySQL®
    • • 14.2 Key–Value Stores
    • • 14.3 Wide-Column Stores
    • • 14.4 Document Stores
    • • 14.5 Further Reading
    • • 14.6 Glossary
  16. Chapter 15: Software Engineering Best Practices

    • • 15.1 Coding Style
    • • 15.2 Version Control and Git for Data Scientists
    • • 15.3 Testing Code
    • • 15.4 Test-Driven Development
    • • 15.5 AGILE Methodology
    • • 15.6 Further Reading
    • • 15.7 Glossary
  17. Chapter 16: Traditional Natural Language Processing

    • • 16.1 Do I Even Need NLP?
    • • 16.2 The Great Divide: Language Versus Statistics
    • • 16.3 Example: Sentiment Analysis on Stock Market Articles
    • • 16.4 Software and Datasets
    • • 16.5 Tokenization
    • • 16.6 Central Concept: Bag-of-Words
    • • 16.7 Word Weighting: TF-IDF
    • • 16.8 n-Grams
    • • 16.9 Stop Words
    • • 16.10 Lemmatization and Stemming
    • • 16.11 Synonyms
    • • 16.12 Part of Speech Tagging
    • • 16.13 Common Problems
    • • 16.14 Advanced Linguistic NLP: Syntax Trees, Knowledge, and Understanding
    • • 16.15 Further Reading
    • • 16.16 Glossary
  18. Chapter 17: Time Series Analysis

    • • 17.1 Example: Predicting Wikipedia Page Views
    • • 17.2 A Typical Workflow
    • • 17.3 Time Series Versus Time-Stamped Events
    • • 17.4 Resampling and Interpolation
    • • 17.5 Smoothing Signals
    • • 17.6 Logarithms and Other Transformations
    • • 17.7 Trends and Periodicity
    • • 17.8 Windowing
    • • 17.9 Brainstorming Simple Features
    • • 17.10 Better Features: Time Series as Vectors
    • • 17.11 Fourier Analysis: Sometimes a Magic Bullet
    • • 17.12 Time Series in Context: The Whole Suite of Features
    • • 17.13 Further Reading
    • • 17.14 Glossary
  19. Chapter 18: Probability

    • • 18.1 Flipping Coins: Bernoulli Random Variables
    • • 18.2 Throwing Darts: Uniform Random Variables
    • • 18.3 The Uniform Distribution and Pseudorandom Numbers
    • • 18.4 Nondiscrete, Noncontinuous Random Variables
    • • 18.5 Notation, Expectations, and Standard Deviation
    • • 18.6 Dependence, Marginal, and Conditional Probability
    • • 18.7 Understanding the Tails
    • • 18.8 Binomial Distribution
    • • 18.9 Poisson Distribution
    • • 18.10 Normal Distribution
    • • 18.11 Multivariate Gaussian
    • • 18.12 Exponential Distribution
    • • 18.13 Log-Normal Distribution
    • • 18.14 Entropy
    • • 18.15 Further Reading
    • • 18.16 Glossary
  20. Chapter 19: Statistics

    • • 19.1 Statistics in Perspective
    • • 19.2 Bayesian Versus Frequentist: Practical Tradeoffs and Differing Philosophies
    • • 19.3 Hypothesis Testing: Key Idea and Example
    • • 19.4 Multiple Hypothesis Testing
    • • 19.5 Parameter Estimation
    • • 19.6 Hypothesis Testing: t-Test
    • • 19.7 Confidence Intervals
    • • 19.8 Bayesian Statistics
    • • 19.9 Naive Bayesian Statistics
    • • 19.10 Bayesian Networks
    • • 19.11 Choosing Priors: Maximum Entropy or Domain Knowledge
    • • 19.12 Further Reading
    • • 19.13 Glossary
  21. Chapter 20: Programming Language Concepts

    • • 20.1 Programming Paradigms
    • • 20.2 Compilation and Interpretation
    • • 20.3 Type Systems
    • • 20.4 Further Reading
    • • 20.5 Glossary
  22. Chapter 21: Performance and Computer Memory

    • • 21.1 A Word of Caution
    • • 21.2 Example Script
    • • 21.3 Algorithm Performance and Big-O Notation
    • • 21.4 Some Classic Problems: Sorting a List and Binary Search
    • • 21.5 Amortized Performance and Average Performance
    • • 21.6 Two Principles: Reducing Overhead and Managing Memory
    • • 21.7 Performance Tip: Use Numerical Libraries When Applicable
    • • 21.8 Performance Tip: Delete Large Structures You Don’t Need
    • • 21.9 Performance Tip: Use Built-In Functions When Possible
    • • 21.10 Performance Tip: Avoid Superfluous Function Calls
    • • 21.11 Performance Tip: Avoid Creating Large New Objects
    • • 21.12 Further Reading
    • • 21.13 Glossary
  23. Chapter 22: Computer Memory and Data Structures

    • • 22.1 Virtual Memory, the Stack, and the Heap
    • • 22.2 Example C Program
    • • 22.3 Data Types and Arrays in Memory
    • • 22.4 Structs
    • • 22.5 Pointers, the Stack, and the Heap
    • • 22.6 Key Data Structures
    • • 22.7 Further Reading
    • • 22.8 Glossary
  24. Chapter 23: Maximum-Likelihood Estimation and Optimization

    • • 23.1 Maximum-Likelihood Estimation
    • • 23.2 A Simple Example: Fitting a Line
    • • 23.3 Another Exam

Customer Reviews

0.0

0 reviews

5 stars
0
4 stars
0
3 stars
0
2 stars
0
1 stars
0

No reviews yet. Be the first to review this book!

Write a Review

Select rating

0/20 characters minimum

By submitting a review, you agree that it may be published after moderation.

Reviewed by GradeFocus Editorial Team

▶Research Sources (14)
  • The Data Science Handbook, eBook by Field Cady | 9781394234509
  • The Data Science Handbook eBook by Field Cady - EPUB | Rakuten ...
  • The Data Science Handbook - Porrúa
  • The data science handbook
  • The data science handbook / Field Cady - Penn State University ...
  • https://lernerbooks.com/products/search_results?se...
  • Sourcebooks, LLC.
  • Please verify you are human - Captcha
  • Python Data Science Handbook
  • Applied Data Sciences - Library Guides - Penn State
  • Practical Data Science with Python or ...
  • The Data Science Handbook (Hardcover) by Field Cady
  • Python Data Science Handbook - 2nd Edition by Jake ...
  • Python Data Science Handbook, 2nd Edition (Second ...

Related Books

Mathematical Modeling for Big Data Analytics

Mathematical Modeling for Big Data Analytics

Passent El-Kafrawy

Multimodal Learning Using Heterogeneous Data

Multimodal Learning Using Heterogeneous Data

Saeid Eslamian

Python in Excel Step-by-Step

Python in Excel Step-by-Step

David Langer

Federated Learning

Federated Learning

Anwesha Mukherjee

Designing the AI-Driven Data Foundations

Designing the AI-Driven Data Foundations

Sanjeev Mohan