Characters & Roles
YellowLine NYC engagement — MHP consulting team
Use these names consistently in modules, exercises, and discussion.
| Character | Role | Narrative function |
|---|---|---|
| Marcus Chen | Operations Manager, YellowLine NYC | Client sponsor; states business pain and constraints |
| Elena Vasquez | Data Architect, MHP | Designs medallion architecture; leads tool evaluation |
| Bob Muller | Junior Data Engineer, MHP | Hands-on builder; trainee proxy |
| Sofia Alvarez | Senior Data Engineer, MHP | Mentors Bob on Databricks / PySpark |
| Priya Sharma | BI Analyst, MHP | Defines KPI questions; builds Power BI dashboard in the story |
| James Okonkwo | Data Analyst, MHP | Validates KPI logic with SQL |
Character Details
Marcus Chen — Operations Manager at YellowLine NYC. Runs a fleet of 400+ yellow taxis across all five boroughs (Manhattan-heavy volume). Frustrated by spreadsheets that break and dashboards that lag behind reality. His three non-negotiable constraints: cost (TCO must be justifiable to the board), performance (dispatch needs near-real-time visibility), and compliance (audit-proof numbers for Q3 regulatory review). He doesn’t care about technology — he cares about outcomes.
Elena Vasquez — Data Architect at MHP. 12 years of experience across banking, logistics, and mobility. Designed the medallion architecture that all three pipelines follow. Vendor-neutral by principle — she evaluates tools on merit, not marketing. Her closing line in Module 7: “Technology is a decision. Architecture is responsibility.” She approves all tool pivots; Bob never “picks vendors” alone.
Bob Muller — Junior Data Engineer at MHP, 2 years out of university. The trainee proxy — when trainees build pipelines, they are Bob. He asks the questions trainees are thinking but might not voice. Elena mentors him through architecture decisions; Sofia guides him through Databricks specifics.
Sofia Alvarez — Senior Data Engineer at MHP. Databricks specialist who guides Bob through PySpark, Delta Lake, and cluster management in the Module 2 lab. Snowflake rebuild is Elena + Bob — Sofia does not own the Snowflake port. Also supports optional streaming and ML stretch labs.
Priya Sharma — BI Analyst at MHP. Captures Marcus’s five KPI questions; Bob materialises twelve Gold kpi_* tables so each question has the right grain. Builds the five-page Power BI dashboard in the story. Her Gold schema is the contract — same tables, same columns, regardless of which pipeline engine produced them.
James Okonkwo — Data Analyst at MHP. Validates KPI logic by writing independent SQL checks against Gold tables — especially kpi_data_quality_metrics in Module 4. Ensures the numbers Priya’s dashboard shows match the source data.
Think & Discuss
Situation (Story day — before Module 1): YellowLine NYC has millions of taxi trips but no analytics platform. Marcus hired MHP. Priya needs KPIs; Elena wants medallion architecture — nothing is built yet. Sketch your initial design; Module 1 formalizes layer semantics. Save this whiteboard — revisit in Module 7.
Prompts:
- What is Marcus’s biggest problem in your own words?
- Where does trip data live today? How often should it be refreshed?
- If you split data into layers, how many would you use and what goes in each?
- Name two tools you would consider for the pipeline. Why those two?
- Priya needs a dashboard — what must exist in the pipeline before she can build it?
- What could go wrong (data quality, cost, skills, maintenance)?
Do not reveal answers yet — Module 1 and labs validate your ideas.
Trainee instruction
Today you are Bob. Elena designed the architecture; your job is to build, evaluate, and recommend tools for Marcus.
Role boundaries (credibility)
- Elena approves tool pivots (Databricks → Snowflake → dbt) — Bob does not “pick vendors” alone.
- Sofia mentors Databricks/PySpark — she does not replace Elena’s architecture sign-off or James’s SQL validation.
- James independently checks Gold KPIs in SQL — he does not redefine Priya’s dashboard layout.
- dbt is always framed as a transform layer on Snowflake, not a warehouse replacement.
- Priya consumes Gold KPI tables in Power BI — same schema regardless of pipeline engine.