Product analytics
From SDK events to dashboards and models
Sources
Ingest
Process
Store
Consume
Web app
Browser SDK
Mobile
iOS and Android
Edge API
Collector
Consent gate
Policy filter
Event stream
Kafka topic
PII vault
Encrypted
Warehouse
Analytics tables
Feature store
Daily batch
Dashboards
Product metrics
ML model
Ranking job
clickstream
app events
identity
accepted events
identity map
normalized facts
daily aggregates
metrics SQL
feature vectors
restricted join
diagram.html#The link keeps the view, selection, route and playback.
About this diagram
Data moving through stages: pipelines, ETL and ELT, change data capture, event streams, lineage.
Ask for one like it
Show how data flows through our feature platform, from the product databases to the models that read it.
The JSON
{ "kind": "dataflow", "density": "compact", "title": "Product analytics", "subtitle": "From SDK events to dashboards and models", "direction": "RIGHT", "phases": [ { "id": "sources", "label": "Sources", "nodes": ["web", "mobile"] }, { "id": "ingest", "label": "Ingest", "nodes": ["edge"] }, { "id": "process", "label": "Process", "nodes": ["consent", "stream"] }, { "id": "store", "label": "Store", "nodes": ["pii", "warehouse", "features"] }, { "id": "consume", "label": "Consume", "nodes": ["dashboards", "model"] } ], "nodes": [ { "id": "web", "type": "client", "card": { "title": "Web app", "subtitle": "Browser SDK" } }, { "id": "mobile", "type": "client", "card": { "title": "Mobile", "subtitle": "iOS and Android" } }, { "id": "edge", "type": "gateway", "card": { "title": "Edge API", "subtitle": "Collector" } }, { "id": "consent", "type": "security", "card": { "title": "Consent gate", "subtitle": "Policy filter" } }, { "id": "stream", "type": "queue", "card": { "title": "Event stream", "subtitle": "Kafka topic" } }, { "id": "pii", "type": "security", "card": { "title": "PII vault", "subtitle": "Encrypted" } }, { "id": "warehouse", "type": "database", "card": { "title": "Warehouse", "subtitle": "Analytics tables", "brand": "snowflake" } }, { "id": "features", "type": "database", "card": { "title": "Feature store", "subtitle": "Daily batch" } }, { "id": "dashboards", "type": "service", "card": { "title": "Dashboards", "subtitle": "Product metrics" } }, { "id": "model", "type": "service", "card": { "title": "ML model", "subtitle": "Ranking job" } } ], "edges": [ { "id": "e1", "from": "web", "to": "edge", "label": "clickstream", "tone": "main" }, { "id": "e2", "from": "mobile", "to": "edge", "label": "app events" }, { "id": "e3", "from": "edge", "to": "consent", "label": "identity", "tone": "security" }, { "id": "e4", "from": "edge", "to": "stream", "label": "accepted events", "tone": "main" }, { "id": "e5", "from": "consent", "to": "pii", "label": "identity map", "tone": "security" }, { "id": "e6", "from": "stream", "to": "warehouse", "label": "normalized facts", "tone": "main" }, { "id": "e7", "from": "warehouse", "to": "features", "label": "daily aggregates", "kind": "async" }, { "id": "e8", "from": "warehouse", "to": "dashboards", "label": "metrics SQL" }, { "id": "e9", "from": "features", "to": "model", "label": "feature vectors", "kind": "async" }, { "id": "e10", "from": "pii", "to": "dashboards", "label": "restricted join", "tone": "security" } ], "notes": [ { "title": "Sensitive boundary", "items": [ "Consent and PII paths are security flows", "PII lands in a restricted vault, apart from the warehouse", "Restricted joins are visible without implying default access" ] } ]}More dataflow examples
All examples- stackmap: JSON to HTMLDataflow9 nodesWritten by an agent“Make an architecture diagram of this repository: what the packages are, what each one does, and how data moves between them from an agent-written diagram JSON to the HTML file a user opens. Back it with evidence from the code.”
- Order event-stream topologyDataflow12 nodes
- ML feature platformDataflow13 nodesWritten by an agent“Show how data flows through our ML feature platform. Product databases (Postgres) are captured with Debezium CDC into Kafka. A Flink job computes streaming features and writes them to Redis (online store) and to a Delta Lake on S3 (offline store). Airflow runs nightly Spark jobs over the lake to build training sets, which a training pipeline on SageMaker consumes; models are registered in MLflow. The prediction service reads online features from Redis and loads models from MLflow. Also note that raw Kafka topics are archived to S3 for replay.”