Overview
- Explains techniques for integrating frameworks for high model performance
- Presents a hybrid approach for rapid prototyping models, deploying and scaling them
- Bridges the gap between machine and deep learning frameworks
Access this book
Tax calculation will be finalised at checkout
Other ways to access
Table of contents (10 chapters)
Keywords
About this book
The book covers an in-memory, distributed cluster computing framework known as PySpark, machine learning framework platforms known as scikit-learn, PySpark MLlib, H2O, and XGBoost, and a deep learning (DL) framework known as Keras.
The book starts off presenting supervised and unsupervised ML and DL models, and then it examines big data frameworks along with ML and DL frameworks. Author Tshepo Chris Nokeri considers a parametric model known as the Generalized Linear Model and a survival regression model known as the Cox Proportional Hazards model along with Accelerated Failure Time (AFT). Also presented is a binary classification model (logistic regression) and an ensemble model (Gradient Boosted Trees). The book introduces DL and an artificial neural network known as the Multilayer Perceptron (MLP) classifier. A way of performing cluster analysis using the K-Means model is covered. Dimension reduction techniques such as Principal Components Analysis and Linear Discriminant Analysis are explored. And automated machine learning is unpacked.
This book is for intermediate-level data scientists and machine learning engineers who want to learn how to apply key big data frameworks and ML and DL frameworks. You will need prior knowledge of the basics of statistics, Python programming, probability theories, and predictive analytics.
What You Will Learn
- Understand widespread supervised and unsupervised learning, including key dimension reduction techniques
- Know the big data analytics layers such as data visualization, advanced statistics, predictive analytics, machine learning, and deep learning
- Integrate big data frameworks with a hybrid of machine learning frameworks and deep learning frameworks
- Design, build, test, and validate skilled machine models and deep learning models
- Optimize model performance using data transformation, regularization, outlier remedying, hyperparameter optimization, and data split ratio alteration
Who This Book Is For
Data scientists and machine learning engineers with basic knowledge and understanding of Python programming, probability theories, and predictive analytics
Reviews
Authors and Affiliations
About the author
Bibliographic Information
Book Title: Data Science Solutions with Python
Book Subtitle: Fast and Scalable Models Using Keras, PySpark MLlib, H2O, XGBoost, and Scikit-Learn
Authors: Tshepo Chris Nokeri
DOI: https://doi.org/10.1007/978-1-4842-7762-1
Publisher: Apress Berkeley, CA
eBook Packages: Professional and Applied Computing, Apress Access Books, Professional and Applied Computing (R0)
Copyright Information: Tshepo Chris Nokeri 2022
Softcover ISBN: 978-1-4842-7761-4Published: 26 October 2021
eBook ISBN: 978-1-4842-7762-1Published: 25 October 2021
Edition Number: 1
Number of Pages: XVI, 119
Number of Illustrations: 35 b/w illustrations
Topics: Statistics, general, Machine Learning, Artificial Intelligence, Python