

Getting Started with Greenplum for Big Data Analytics. A hands-on guide on how to execute an analytics project from conceptualization to operationaliz


Getting Started with Greenplum for Big Data Analytics. A hands-on guide on how to execute an analytics project from conceptualization to operationaliz - Najlepsze oferty
Getting Started with Greenplum for Big Data Analytics. A hands-on guide on how to execute an analytics project from conceptualization to operationaliz - Opis
Organizations are leveraging the use of data and analytics to gain a competitive advantage over their opposition. Therefore, organizations are quickly becoming more and more data driven. With the advent of Big Data, existing Data Warehousing and Business Intelligence solutions are becoming obsolete, and a requisite for new agile platforms consisting of all the aspects of Big Data has become inevitable. From loading/integrating data to presenting analytical visualizations and reports, the new Big Data platforms like Greenplum do it all. It is now the mindset of the user that requires a tuning to put the solutions to work.Getting Started with Greenplum for Big Data Analytics is a practical, hands-on guide to learning and implementing Big Data Analytics using the Greenplum Integrated Analytics Platform. From processing structured and unstructured data to presenting the results/insights to key business stakeholders, this book explains it all.Getting Started with Greenplum for Big Data Analytics discusses the key characteristics of Big Data and its impact on current Data Warehousing platforms. It will take you through the standard Data Science project lifecycle and will lay down the key requirements for an integrated analytics platform. It then explores the various software and appliance components of Greenplum and discusses the relevance of each component at every level in the Data Science lifecycle.You will also learn Big Data architectural patterns and recap some key advanced analytics techniques in detail. The book will also take a look at programming with R and (...) więcej integration with Greenplum for implementing analytics. Additionally, you will explore MADlib and advanced SQL techniques in Greenplum for analytics. This book also elaborates on the physical architecture aspects of Greenplum with guidance on handling high-availability, back-up, and recovery. Spis treści:Getting Started with Greenplum for Big Data Analytics
Table of Contents
Getting Started with Greenplum for Big Data Analytics
Credits
Foreword
About the Author
Acknowledgement
About the Reviewers
www.PacktPub.com
Support files, eBooks, discount offers and more
Why Subscribe?
Free Access for Packt account holders
Instant Updates on New Packt Books
Preface
What this book covers
What you need for this book
Who this book is for
Conventions
Reader feedback
Customer support
Errata
Piracy
Questions
1. Big Data, Analytics, and Data Science Life Cycle
Enterprise data
Classification
Features
Big Data
So, what is Big Data?
Multi-structured data
Data analytics
Data science
Data science life cycle
Phase 1 state business problem
Phase 2 set up data
Phase 3 explore/transform data
Phase 4 model
Phase 5 publish insights
Phase 6 measure effectiveness
References/Further reading
Summary
2. Greenplum Unified Analytics Platform (UAP)
Big Data analytics platform requirements
Greenplum Unified Analytics Platform (UAP)
Core components
Greenplum Database
Hadoop (HD)
Chorus
Command Center
Modules
Database modules
HD modules
Data Integration Accelerator (DIA) modules
Core architecture concepts
Data warehousing
Column-oriented databases
Parallel versus distributed computing/processing
Shared nothing, massive parallel processing (MPP) systems, and elastic scalability
Shared disk data architecture
Shared memory data architecture
Shared nothing data architecture
Data loading patterns
Greenplum UAP components
Greenplum Database
The Greenplum Database physical architecture
The Greenplum high-availability architecture
High-speed data loading using external tables
External table types
Polymorphic data storage and historic data management
Data distribution
Hadoop (HD)
Hadoop Distributed File System (HDFS)
Hadoop MapReduce
Chorus
Greenplum Data Computing Appliance (DCA)
Greenplum Data Integration Accelerator (DIA)
References/Further reading
Summary
3. Advanced Analytics Paradigms, Tools, and Techniques
Analytic paradigms
Descriptive analytics
Predictive analytics
Prescriptive analytics
Analytics classified
Classification
Forecasting or prediction or regression
Clustering
Optimization
Simulations
Modeling methods
Decision trees
Association rules
The Apriori algorithm
Linear regression
Logistic regression
The Naive Bayesian classifier
K-means clustering
Text analysis
R programming
Weka
In-database analytics using MADlib
References/Further reading
Summary
4. Implementing Analytics with Greenplum UAP
Data loading for Greenplum Database and HD
Greenplum data loading options
External tables
gpfdist
gpload
Hadoop (HD) data loading options
Sqoop 2
Greenplum BulkLoader for Hadoop
Using external ETL to load data into Greenplum
Extraction, Load, and Transformation (ELT) and Extraction, Transformation, Load, and Transformation (ETLT)
Greenplum target configuration
Sourcing large volumes of data from Greenplum
Unsupported Greenplum data types
Push Down Optimization (PDO)
Greenplum table distribution and partitioning
Distribution
Data skew and performance
Optimizing the broadcast or redistribution motion for data co-location
Partitioning
Querying Greenplum Database and HD
Querying Greenplum Database
Analyzing and optimizing queries
The ANALYZE function
The EXPLAIN function
Dynamic Pipelining in Greenplum
Querying HDFS
Hive
Pig
Data communication between Greenplum Database and Hadoop (using external tables)
Data Computing Appliance (DCA)
Storage design, disk protection, and fault tolerance
Master server RAID configurations
Segment server RAID configurations
Monitoring DCA
Greenplum Database management
In-database analytics options (Greenplum-specific)
Window functions
The PARTITION BY clause
The ORDER BY clause
The OVER (ORDER BY) clause
Creating, modifying, and dropping functions
User-defined aggregates
Using R with Greenplum
DBI Connector for R
PL/R
Using Weka with Greenplum
Using MADlib with Greenplum
Using Greenplum Chorus
Pivotal
References/Further reading
Summary
Index O autorze: Sunila Gollapudi works as Vice President Technology with Broadridge Financial Solutions (India) Pvt. Ltd., a wholly owned subsidiary of the US-based Broadridge Financial Solutions Inc. (BR). She has close to 14 years of rich hands-on experience in the IT services space. She currently runs the Architecture Center of Excellence from India and plays a key role in the big data and data science initiatives. Prior to joining Broadridge she held key positions at leading global organizations and specializes in Java, distributed architecture, big data technologies, advanced analytics, Machine learning, semantic technologies, and data integration tools. Sunila represents Broadridge in global technology leadership and innovation forums, the most recent being at IEEE for her work on semantic technologies and its role in business data lakes. Sunila's signature strength is her ability to stay connected with ever changing global technology landscape where new technologies mushroom rapidly, connect the dots and architect practical solutions for business delivery. A post graduate in computer science, her first publication was on Big Data Datawarehouse solution, Greenplum titled Getting Started with Greenplum for Big Data Analytics, Packt Publishing. She's a noted Indian classical dancer at both national and international levels, a painting artist, in addition to being a mother, and a wife. mniej
Getting Started with Greenplum for Big Data Analytics. A hands-on guide on how to execute an analytics project from conceptualization to operationaliz - Opinie i recenzje
Na liście znajdują się opinie, które zostały zweryfikowane (potwierdzone zakupem) i oznaczone są one zielonym znakiem Zaufanych Opinii. Opinie niezweryfikowane nie posiadają wskazanego oznaczenia.