Serverless ETL and Analytics with AWS Glue - Second Edition

Serverless ETL and Analytics with AWS Glue - Second Edition

Albert Quiroga / Noritaka Sekiyama / Tomohiro Tanaka

71,26 €
IVA incluido
Disponible
Editorial:
Packt Publishing
Año de edición:
2026
ISBN:
9781835464847
71,26 €
IVA incluido
Disponible

Selecciona una librería:

  • Librería Samer Atenea
  • Kálamo Books
  • Librería Elías (Asturias)
  • Librería Kolima (Madrid)
  • Librería Proteo (Málaga)

Use AWS Glue to integrate growing data sources with serverless ETL, building secure, observable pipelines that support reliable analytics while managing performance and cost across a governed AWS data platform as workloads growKey Features:- Use runnable code, console walkthroughs, and downloadable examples for core AWS Glue workflows- Apply DataOps practices with AWS CDK, Docker, and CI/CD in real-world scenarios- Learn from six data specialists with AWS, Spark, Apache Iceberg, and data lake expertiseBook Description:Whether you build data pipelines, design cloud architectures, or deliver analytics on AWS, bringing data together is only part of the challenge. You must also keep this data clean, trustworthy, and available while controlling costs. AWS Glue offers serverless data integration, but using it effectively requires decisions about storage, metadata, security, orchestration, monitoring, and performance.This book guides you from modern data management and core AWS Glue features through ingestion from files, streams, SaaS applications, and JDBC sources, preparation, storage layout, metadata, security, sharing, and pipeline operations. Console walkthroughs and runnable examples show how to manage schemas and lineage in AWS Glue Data Catalog, apply AWS Lake Formation access controls, monitor workloads, tune Spark jobs, troubleshoot failures, and manage development with AWS CDK, Docker, and CI/CD. You will also examine analytics, machine learning and generative AI integrations, real-world data lake scenarios, and cost optimization. Learn how Apache Iceberg, Apache Hudi, and Delta Lake add transactions, schema evolution, and efficient data management to data lakes.By the end, you will be able to design, build, operate, and continuously improve a serverless data platform that fits your organization’s scale, structure, and priorities.What You Will Learn:- Design scalable serverless ETL pipelines with AWS Glue- Ingest data from files, streams, SaaS, and JDBC sources- Optimize file formats, partitions, compression, and layouts- Manage schemas, partitions, and lineage in AWS Glue Data Catalog- Secure data with access control, encryption, and auditing- Automate testing and multi account CI/CD using AWS CDK and Docker- Monitor, tune, and troubleshoot AWS Glue and Spark workloads- Apply Apache Iceberg, Hudi, and Delta Lake to data lakes with AWS GlueWho this book is for:This book is for data engineers, ETL developers, cloud architects, and analytics professionals who build or operate data platforms on AWS. It suits readers working on serverless data lakes, Spark ETL, governance, data sharing, reliability, or cost control. It is especially useful if you aim to improve pipeline reliability, governance, or cost visibility as workloads grow. Basic familiarity with the AWS Management Console, Amazon S3, and IAM is recommended. Experience with Python, SQL, or Apache Spark will help with the code examples, and an AWS account is useful for following the walkthroughs.Table of Contents- Data Management: Introduction and Concepts- Introduction to Important AWS Glue Features- Data Ingestion- Data Preparation- Data Layouts- Data Management- Implementing Resilient Metadata Management- Data Security- Data Sharing- Data Pipeline Management- Monitoring- Tuning, Debugging and Troubleshooting- Data Analysis- Machine Learning and Generative AI Integration- Architecting Data Lakes for Real World Scenarious and Edge Cases- End-to-end development lifecycle- Open Table Format- Cost Optimization

Artículos relacionados

  • Exploring Advances in Interdisciplinary Data Mining and Analytics
    Data mining is still a relatively young field, expanding at the rate of technology while advancing tools and techniques for gaining knowledge, finding patterns, and managing databases. Exploring Advances in Interdisciplinary Data Mining and Analytics: New Trends is an updated look at the state of technology in the field of data mining and analytics. As processor speeds, databas...
  • Knowledge Discovery Practices and Emerging Applications of Data Mining
    Recent developments have drastically increased the volume and complexity of data available to be mined, leading researchers to explore new ways to glean non-trivial data automatically. Knowledge Discovery Practices and Emerging Applications of Data Mining: Trends and New Domains introduces the reader to recent research activities in the field of data mining. This book covers as...
  • Research and Trends in Data Mining Technologies and Applications
    David Taniar
    ...
  • Developing Metadata Application Profiles
    The prevalence of data science has grown exponentially in recent years. Increases in data exchange have created the need for standards and formats on handling data from different sources. Developing Metadata Application Profiles is an innovative reference source that discusses the latest trends and techniques for effectively managing and exchanging metadata. Including a range o...
  • Modern Technologies for Big Data Classification and Clustering
    Data has increased due to the growing use of web applications and communication devices. It is necessary to develop new techniques of managing data in order to ensure adequate usage. Modern Technologies for Big Data Classification and Clustering is an essential reference source for the latest scholarly research on handling large data sets with conventional data mining and provi...
  • 90 Gelöste Fälle zu Zeitintelligenz in der DAX-Sprache
    Ramón Javier Castro Amador
    Dieser Ratgeber ist rein praktisch ausgerichtet, so dass Sie den gesamten DAX-Code in dieser Publikation anhand einer zum Download verfügbaren .pbix-Datei testen können.'90 gelöste Fälle zu Zeitintelligenz in DAX' ist ein Ratgeber für Benutzer von Microsoft Power BI, der Lösungen für sehr häufige praktische Fälle in Zeitintelligenzmodellen in der Sprache DAX bietet.Um das Verst...
    Disponible

    16,15 €