Databricks Trainer-Dutch

Il y a 6 jours

Brussels, Brussels-Capital, Belgique Phoenix Technologies Temps plein
We would like the training to cover the following topics:
• Databricks Functionality
• Databricks Unity Catalog (specifically including the Lineage View)
• Using Notebooks – %md | %sql | %python
• The MS Word-like look and feel of Notebooks
• Notebooks Revision History functionality
• Notebooks Export/Import functionality
• Visualization capabilities in Databricks (both visuals within a cell and separate dashboards)
• Using Genie
• Defining temporary views to build upon in subsequent commands
• Reading Excel files and writing them to a temporary view
• Interaction with Power BI – How to load Databricks output datasets into Power BI
• Delta Lake storage and partitioning
• Performance analysis – identifying the steps the cluster executes and pinpointing where slowdowns occur o Useful PySpark libraries ▪ Installing/importing libraries ▪ PySpark DataFrame and SQL library (pyspark.sql) and scenarios where this is preferred over SQL
• Advanced SQL: Nested queries, JOIN clause conditions, CASE WHEN, querying metadata (Information Schema), VERSION AS OF
• Using parameters in Databricks
• Scheduling workflows in Databricks
• Serverless vs. Dedicated Clusters -> Which to use when?
• … Internal
• Client Data and Information Platform:
• Architecture diagram (bronze-silver-gold)
• Domain catalogs and recurring schemas
• Figures regarding current usage within Client and/or the number of tables on Databricks
• Data Vault Modeling – ref. our BDV (Business Data Vault) schema in the silver layer
• Dimensional Modeling – ref. our DWH (Data Warehouse) schema in the gold layer
• Consulting the Repos logic: 'federated' code that feeds the various layers (bronze-silver-gold) on the platform
• Where to store self-service tables and logic (analysis_ungoverned and playground schemas for data, and folders for logic) We expect a significant number of exercises and case studies relevant to Client to be developed during the training, based on Client's internal data. Learning objectives: Upon completion of this training, participants will be able to:
• Independently build self-service ETL processes on the Databricks platform using SQL and/or PySpark
• Make optimal use of Databricks features.
• To design a logical data model that can be technically implemented by a more specialized Data Engineer
• To be able to develop a prototype star schema that can be technically implemented by a more specialized Data Engineer
• To schedule workflows and optimize queries Skills Required In the case of a classroom training, the teacher can be present at the requested location at 7.30 am at the latest The teacher can make himself available within 6 weeks to give a training (from request for new session(s)) The trainer has demonstrable experience as an instructor The trainer has a thorough knowledge of and experience with the Databricks Platform The trainer has a strong command of programming languages SQL and Python for data analysis and machine learning on Databricks The trainer can move to all specified locations The trainer shows in the plan of action that he has provided similar training courses in the past. We expect at least 2 reference examples (where, when, content/short description, sector/customer,.. The trainer has a certification from Databricks The trainer has experience in implementing advanced data warehousing and data governance processes The trainer has experience in developing and maintaining machine learning models and data visualization tools. The trainer has at least 2 years of experience in the energy sector NOT Mandatory Its plus