Get Free Shipping on orders over $79
Mining Imperfect Data : With Examples in R and Python - Ronald K. Pearson

Mining Imperfect Data

With Examples in R and Python

By: Ronald K. Pearson

Paperback | 28 February 2021 | Edition Number 2

At a Glance

Paperback


RRP $224.00

$211.75

or 4 interest-free payments of $52.94 with

 or 

Ships in 5 to 7 business days

It has been estimated that as much as 80% of the total effort in a typical data analysis project is taken up with data preparation, including reconciling and merging data from different sources, identifying and interpreting various data anomalies, and selecting and implementing appropriate treatment strategies for the anomalies that are found. This book focuses on the identification and treatment of data anomalies, including examples that highlight different types of anomalies, their potential consequences if left undetected and untreated, and options for dealing with them.

As both data sources and free, open-source data analysis software environments proliferate, more people and organizations are motivated to extract useful insights and information from data of many different kinds (e.g., numerical, categorical, and text). The book emphasizes the range of open-source tools available for identifying and treating data anomalies, mostly in R but also with several examples in Python.

Mining Imperfect Data: With Examples in R and Python, Second Edition
  • presents a unified coverage of 10 different types of data anomalies (outliers, missing data, inliers, metadata errors, misalignment errors, thin levels in categorical variables, noninformative variables, duplicated records, coarsening of numerical data, and target leakage);
  • includes an in-depth treatment of time-series outliers and simple nonlinear digital filtering strategies for dealing with them; and
  • provides a detailed introduction to several useful mathematical characteristics of important data characterizations that do not appear to be widely known among practitioners, such as functional equations and key inequalities.

More in Information Technology General Issue

Careless People : A story of where I used to work - Sarah Wynn-Williams

RRP $24.99

$21.75

13%
OFF
Doppelganger : A Trip Into the Mirror World - Naomi Klein

RRP $26.99

$22.99

15%
OFF
Ethics, Information, and Technology : A Tangled Web - Kip Currier

RRP $110.00

$96.75

12%
OFF
Against the Machine : On the Unmaking of Humanity - Paul Kingsnorth
Gilded Rage : Elon Musk and the Radicalization of Silicon Valley - Jacob Silverman
Building a Scalable Data Warehouse with Data Vault 2.0 - Dan Linstedt
Power On : Managing Screen Time to Benefit the Whole Family - Ash Brandin
Developing Graphics Frameworks with Java and OpenGL - Lee Stemkoski
Computer Organization and Architecture, Global Edition : 11th Edition - William Stallings
Business Driven Information Systems ISE : 9th Edition - Paige Baltzan