Senior Data Engineer (Pipelines & Architecture)
Massive is an infrastructure company powering data-driven innovation at scale. We operate a global proxy network, build managed APIs for scraping, and develop AI-powered automation tools for large-scale data collection and processing.
Our Monetization SDK is a cross-platform development kit that enables app developers to offer users an alternative payment method, opting in to share small amounts of unused computing resources instead of paying with money or viewing ads. The SDK powers millions of users worldwide across desktop, mobile, and smart TV platforms.
Responsibilities
- Take the streaming pipeline from design to production, then operate it at roughly 5 TB/day.
- Move usage aggregation off the legacy database onto the new streaming path, while everything downstream, invoices, dashboards, reports, keeps working through the migration.
- Design the storage layer that decides what the system costs, how events are laid out, partitioned, retained, and expired.
- Make ingestion correct under failure: replays, duplicates, out-of-order arrivals, and late data that shows up after a window has already closed.
- Own the aggregation layer that customer-facing usage views and internal reporting read from, the models, the metric definitions, and the freshness guarantees behind them.
- Keep performance predictable at both ends of the system, the ingest firehose and the queries analysts run against it, as traffic grows.
- Automate what is currently manual: schema changes tracked in code, backfills and reconciliation as repeatable jobs, no ad-hoc surgery on production data.
- Instrument the platform for freshness, lag, ingest errors, and cost, with alerts that fire before anyone notices a wrong number on a dashboard.
- Participate in on-call and incident response for the data platform, and drive the follow-ups that keep the same failure from recurring.
- Collaborate closely with cross-functional teams (backend, proxy infrastructure, product) so the events we emit answer the questions the business actually asks.
Requirements
- 7+ years building and operating production data pipelines.
- Hands-on ClickHouse experience: engine selection (ReplacingMergeTree and friends), partitioning, materialized views, and query tuning across billions of rows.
- Strong PostgreSQL: partitioning, indexing, bulk-insert performance, and a clear view of where it stops being the right tool.
- Deep SQL and analytical data modeling, turning raw events into aggregates other people can trust without asking you first.
- Demonstrated ownership, you've been the person accountable for a pipeline that was still correct the morning after it broke.
- Active user of agentic coding tools.
- Strong communication skills, you write clearly, escalate early, and make your reasoning visible in an async team.
- Self-directed and comfortable in a fast-paced startup environment where the scope is yours to define.
Bonus Points
- Kafka, or a comparable streaming platform, run in production, partitioning, consumer groups, and delivery guarantees.
- Go, our pipeline services are written in it, and you'd be writing and debugging them.
- Experience with high-volume telemetry or observability data (multi-TB/day ingest).
- Hands-on with Vector or comparable collectors (Fluent Bit, OpenTelemetry Collector, Logstash).
- Experience with usage-based metering or billing pipelines, including reconciliation and late-data handling.
- Workflow orchestration for backfills and scheduled jobs (Temporal, Airflow, Dagster).
- BI and transformation tooling, Metabase, Grafana, dbt or equivalent.
- AWS experience (RDS, Lambda, SQS, S3/Parquet, EC2) and Terraform.
- Experience self-hosting and operating ClickHouse or Kafka clusters rather than consuming them as a service.