Duration: ~60β90 minutes
Difficulty: Intermediate
Technical Requirements: Working knowledge of cloud platforms (AWS or Azure), SQL, and basic command-line operations
Workshop Type: This workshop is currently designed to work in these three modes:
This hands-on workshop demonstrates a real-time instant-payments operations pipeline for RiverPay, a fictitious mid-size payments processor that sits behind ~40 regional banks and credit unions.
You will ingest customer profiles and FX rates via CDC, stream RiverFlow payment lifecycle events (multi-currency), enrich them with Flink to create three data products β completed payments (with FX temporal join), an operational risk_score (profile temporal join + external risk UDF), and a trailing-24h customer risk exposure aggregate per customer β then sync those products via Tableflow into Databricks Genie (RiverPulse).
RiverPay's partner banks need instant-payments parity, but ops tooling is still batch-based. End-of-day reports cannot answer: which payment needs attention right now? This workshop is an operational-visibility story where risk_score means operational exception probability β not fraud.
- RiverFlow is RiverPay's instant-payments rail (a four-stage lifecycle that maps to FedNow/RTP-style flows).
- RiverPulse is the real-time ops/analytics layer on top β Tableflow into Databricks Genie.
- Which payments are most likely to need manual intervention right now?
- Which customers drive the highest operational exception exposure in the last 24 hours?
- What is the RiverFlow lifecycle completion rate from initiation to completed status? (Phase 1 proxy: completed = 4-way join + FX enrichment; stall drill-down is backlog)
Expand the accordion below for more background. Otherwise, continue to Datasets and Workshop Labs.
If you have any issues with or feedback for this workshop, Please let us know in this quick 2-minute survey!
Use Case Details
As instant-payments volume grows, RiverPay's ops team is flying blind between batch report runs. Partner banks expect FedNow/RTP-style parity; RiverPay needs real-time signal on payments that are stuck or likely to need a human β without building and maintaining custom lakehouse pipelines.
- Dana Ruiz, VP of Payment Operations β owns βwhich payments need manual intervention right now?β
- Marcus Chen, Head of Data Platform β wants governed data in Databricks without custom pipeline toil (Tableflow)
- Priya Anand, Compliance & Risk Lead β audience for a light PII / CSFLE talking point (not a full security lab)
- Capture customer profiles and FX rates from PostgreSQL with CDC
- Stream payment lifecycle events (initiation β authorization β balance update β status)
- Produce Flink data products β completed payments (4-way inner join + FX TTJ), operational
risk_score(profile TTJ + external risk UDF), and trailing-24h customer risk exposure per customer (OVERwindow + primary key = upsert) - Serve those products via Tableflow into Unity Catalog (TTL / right-to-forget talking point)
- Analyze the data with natural language using Databricks Genie
Tableflow publishes only the three Flink data products (riverflow_payments append, riverflow_payments_risk_score append, riverflow_customer_risk_exposure_24h upsert). Raw lifecycle topics stay Kafka sources.
- Infrastructure as Code β Deploy Confluent Cloud + cloud + Databricks resources with Terraform
- Change Data Capture β Stream Postgres profiles and FX rates into Kafka
- Stream Processing β Flink SQL: multi-stream joins, temporal joins, external UDF
- Data Lake Integration β Tableflow to Delta Lake / Unity Catalog with data TTL
- Ops Analytics β Answer RiverPulse questions in Databricks Genie
- Data Sources β ShadowTraffic (profiles, FX, lifecycle) + PostgreSQL + shared Risk Scoring API
- Ingestion β Postgres CDC connector + Kafka producers for lifecycle topics
- Processing β Apache Flink SQL (completed-payments MT + risk-score MT + customer risk exposure MT)
- Integration β Confluent Tableflow β Delta Lake (S3 or ADLS Gen2)
- Analytics β Databricks Unity Catalog + Genie (RiverPulse)
Full narrative skin: USECASE.md. Architecture notes: context/fsi_payments_workshop_architecture.md.
| Dataset | Source | Topic / table | Tableflow |
|---|---|---|---|
| Customer profiles | Postgres β CDC | riverflow.riverpay.customer_profiles |
No (Kafka source) |
| FX rates | Postgres β CDC (upserts ~5s) | riverflow.riverpay.fx_rates |
No |
| Payment initiation | ShadowTraffic β Kafka | riverflow.payments.initiation |
No |
| Authorization | ShadowTraffic β Kafka | riverflow.payments.authorization |
No |
| Balance update | ShadowTraffic β Kafka | riverflow.payments.balance_update |
No |
| Status | ShadowTraffic β Kafka | riverflow.payments.status |
No |
| Completed payments | Flink MT (inner join + FX TTJ) | riverflow_payments |
Yes (append) |
| Risk score | Flink MT (profile TTJ + risk UDF) | riverflow_payments_risk_score |
Yes (append) |
| Customer risk exposure | Flink MT (trailing-24h OVER aggregate) | riverflow_customer_risk_exposure_24h |
Yes (upsert) |
This workshop supports multiple modes. Choose the path that matches your situation:
Warning
Prerequisites and cost
Cloud paths typically need Confluent Cloud (admin API key), a Unity Catalogβenabled Databricks workspace, AWS and/or Azure, Git, and Docker Desktop. terraform apply creates billable resources β plan to run the cleanup lab when you finish.
Hands-on with Confluent Cloud and Databricks products. Cloud accounts and infrastructure are pre-provisioned. You write Flink SQL, enable Tableflow, and answer RiverPulse questions in Genie.
Use it only when instructed by your workshop instructor/leader.
Operators: Shared Azure infra + per-attendee stacks β
docs/operator-azure-elevate.mdandwsa-spec-azure.yaml.
| Lab | Duration | Details |
|---|---|---|
| LAB 1: Claim Your Account | ~5 min | Claim your workshop account: complete the claim form, receive credentials, verify Confluent Cloud and Databricks. |
| LAB 2: Explore Your Environment | ~10 min | Tour your environment: CDC topics, lifecycle topics, Flink compute pool, risk CONNECTION/UDF. |
| LAB 3: Stream Processing | ~20 min | Transform streams: Flink MTs β completed payments (FX temporal join) + operational risk (profile TTJ + UDF). |
| LAB 4: Tableflow | ~10 min | Enable Tableflow: publish the three Flink data products; TTL / right-to-forget talking point. |
| LAB 5: RiverPulse Analytics | ~15 min | Ask Genie: answer the three RiverPulse business questions. |
| LAB 6: Wrap Up | ~5 min | Recap: review accomplishments and next steps. |
Fully hands-on: you sign up for your own cloud accounts, deploy with Terraform, then build Flink MTs and enable Tableflow yourself.
Use it to learn how you could run a similar pipeline in your own Confluent Cloud and Databricks environments.
Terraform:
terraform/aws(recommended) orterraform/azurewith emptyshared_*.
| Lab | Duration | Details |
|---|---|---|
| LAB 0: Prerequisites | ~10 min | Set up prerequisites: cloud accounts, Git, Docker, clone the repo, build images. |
| LAB 1: Account Setup | ~15 min | Configure credentials: Confluent Cloud API keys, Databricks service principal, AWS or Azure auth, terraform.tfvars. |
| LAB 2: Deploy & Explore | ~20β45 min | Deploy with Terraform: provision infra, CDC, lifecycle traffic, Risk API/UDF plumbing; tour the environment. |
| LAB 3: Stream Processing | ~20 min | Transform streams: write Flink MTs for FX-aware completed payments and operational risk. |
| LAB 4: Tableflow | ~10 min | Enable Tableflow: sync Flink data products to Unity Catalog; TTL talking point. |
| LAB 5: RiverPulse Analytics | ~15 min | Ask Genie: answer the three RiverPulse business questions. |
| LAB 6: Wrap-up & Cleanup | ~10 min | Tear down: terraform destroy and recap. |
Builds on self-service by automating almost all pipeline steps (Flink MTs + Tableflow included). Best for short-term, long-term, or always-on demos with minimal in-product setup.
Use it to show immediate value after one
terraform apply.Demo Terraform root:
terraform/aws-demo.
| Lab | Duration | Details |
|---|---|---|
| LAB 0: Prerequisites | ~10 min | Set up prerequisites: Confluent Cloud, Databricks (UC), AWS, Git, Docker image. |
| LAB 1: Account Setup | ~15 min | Configure credentials: API keys, Databricks SP, AWS, terraform.tfvars. |
| LAB 2: Deploy and Observe | ~20β25 min | Deploy everything: one apply provisions AWS, Confluent, Flink MTs, Tableflow, UC; guided pipeline tour. |
| LAB 3: RiverPulse Analytics | ~15 min | Ask Genie: answer the three RiverPulse business questions. |
| LAB 4: Cleanup | ~10 min | Tear down: terraform destroy and leftover checks. |
Parallel RiverPay-lite path on Red Hat OpenShift Service on AWS (ROSA): Confluent for Kubernetes + Confluent Platform, lifecycle topics visible in Control Center. No Flink / Tableflow / Databricks on this path yet.
Terraform:
terraform/cp-rosa/(Stage 1 then Stage 2). Labs:labs/cp-rosa/.
| Lab | Duration | Details |
|---|---|---|
| LAB 0: Prerequisites | ~10 min | Set up prerequisites: accounts, tools, clone. |
| LAB 1: Account Setup | ~15 min | Configure credentials for the ROSA / CP path. |
| LAB 2: Provision ROSA | ~30β45+ min | Stage 1 Terraform: provision ROSA HCP. |
| LAB 3: Deploy and Observe | ~15β25 min | Stage 2: CFK + CP + RiverPay-lite; observe in Control Center. |
| LAB 4: Cleanup | ~20β40 min | Tear down ROSA / CP resources. |
- Recap: Summary of accomplishments and talking points
- Troubleshooting (Cloud): Common issues for demo, self-service, and instructor-led
- Troubleshooting (cp-rosa): ROSA / CP path issues
- Genie prompts: Suggested RiverPulse prompts and expected answer shape
Congratulations β you've completed the RiverPay hands-on workshop on real-time payments ops with Confluent and Databricks!
Important
Your Feedback Helps!
Please help us improve this workshop by leaving your feedback in this quick 2-minute survey!
Thanks!
See LICENSE.
