Hire
Data Systems Scala Developers
Built through years working with teams and engineers in the Scala ecosystem.
We specialize in functional backend systems, data platforms, and distributed infrastructure.
Focused, relevant introductions from a curated network.
Available Developers
Data Engineer
Engineered large-scale data processing pipelines using Apache Spark and Scala, significantly enhancing data throughput. Developed real-time data ingestion systems with Apache Kafka, ensuring high availability and fault tolerance. Expert in streaming analytics, applying Flink for dynamic data transformations in IoT applications.
- Integrated Scala Play for responsive web data visualization
- Utilized Cats Effect for functional programming concurrency control
- Designed a custom schema registry for Kafka message validation
- Implemented end-to-end data encryption in distributed environments
Big Data Engineer
Led the design and implementation of distributed data processing systems handling petabyte-scale datasets. Spearheaded the integration of Kafka with Spark for real-time analytics pipelines in financial services. Architected a high-performance data ingestion framework using Akka to optimize processing speed and reliability.
- Developed custom SQL query optimization techniques for large datasets
- Implemented Git versioning for complex data workflows
- Pioneered Akka-based microservices for scalable data processing
- Optimized Spark jobs reducing runtime by 30%
Senior Data Engineer
- Led migration of monolithic system to microservices architecture
- Built high-throughput data pipeline processing 1M+ events per second
- Designed and implemented real-time monitoring and alerting platform
Spark Performance Engineer - Data Engineer
Designed and optimized large-scale distributed data processing systems using Apache Spark, focusing on performance enhancements for data-intensive applications. Developed custom Spark transformations and actions, significantly reducing execution time for ETL pipelines in the financial analytics domain. Integrated Rust components with Scala to improve computational efficiency in real-time data processing.
- Implemented advanced caching strategies for Spark job optimization.
- Contributed to open-source Apache DataFusion project with Rust.
- Developed PySpark solutions for streaming data pipelines.
- Led migration of legacy data systems to cloud-based infrastructure.
Staff Data Engineer
Architected large-scale data pipelines using Apache Spark and Scala, optimizing for real-time analytics in the e-commerce domain. Spearheaded the migration of data infrastructure to Databricks, enhancing processing speed and reliability. Developed a custom data validation framework in Python, ensuring data integrity across distributed systems.
- Implemented Rust-based data transformation tools for performance gains
- Designed SQL-based ETL workflows for complex data models
- Automated data quality checks using custom-built Python scripts
Staff Protocol Engineer
Designed MLS group-state features for secure messaging, including per-field permissions, in-place migration of live groups, and invite flows based on external commits. Built a Rust and Go bidirectional streaming transport with ordered delivery, heartbeats, reconnect handling, and database transaction boundaries enforced through async traits.
- Reduced 100-device invite payloads from 8 MB to 200 KB
- Cut full group loads from 13 database reads to one
- Built Rust-backed GraphQL and FFI layers for untrusted drivers
- Worked on edge computer-vision and blockchain-indexing systems
Software Engineer / Data Engineer
Developed large-scale distributed data processing systems for real-time analytics. Specialized in building high-performance ETL pipelines and data warehousing solutions. Designed and optimized custom consensus protocols for fault-tolerant systems.
- Implemented complex event-driven architectures for financial services.
- Architected scalable microservices using Go and Rust.
- Enhanced database internals for improved query performance.
Principal Agentic AI & Lead Data Systems Architect
Architected a multi-agent AI system that integrates generative AI with real-time data processing, optimizing decision-making workflows in complex environments. Spearheaded the development of a scalable graph-based reasoning engine utilizing GraphRAG and LangGraph, enhancing the efficiency of large-scale data analysis. Innovated a modular orchestration framework that leverages LangChain for seamless agent communication and coordination across distributed systems.
- Designed VertexAI-powered pipelines for real-time data ingestion.
- Implemented dynamic agent collaboration using custom orchestration protocols.
- Optimized graph traversal algorithms for high-performance data querying.
- Developed cross-platform AI models with robust fault tolerance.
Senior Engineer / Technical Lead
Led the design and implementation of a distributed event streaming platform using Akka and Kafka. Spearheaded big data processing initiatives with Spark on AWS for analytics.
- Developed large-scale data pipelines with Apache Spark
- Managed team of engineers for complex system integrations
- Optimized server-side rendering with Play Framework
Founding Database Engineer
Pioneered database internals and distributed systems design. Advanced knowledge in operations and systems design using AWS.
- Pioneered database internals
- Expert in distributed systems design
- Advanced AWS operations knowledge
Data Engineer / Contractor
Specialized in building data pipelines using Scala and Java for large-scale ETL processes. Designed graph database solutions using Cypher and CQL to improve data retrieval efficiency.
- Automated data processing workflows with Groovy
- Enhanced data integration using Python scripts
- Implemented graph algorithms for advanced data insights
Freelance Software | Data | ML Engineer
Developed machine learning pipelines for real-time data analysis and prediction. Deployed distributed systems on AWS Fargate for scalable processing.
- Optimized text processing with TF-IDF and k-NN
- Automated data migration using AWS DMS
- Implemented dimensionality reduction with PCA for feature optimization
Software Engineer IV
Developed high-performance distributed systems for financial trading platforms, optimizing latency and throughput for real-time data processing. Designed and implemented scalable consensus protocols for blockchain networks, enhancing security and reliability. Spearheaded the migration of legacy systems to cloud-native architectures, significantly improving deployment efficiency and scalability.
- Architected microservices for telecom billing systems handling millions of transactions
- Built custom compilers for domain-specific languages in the automotive industry
- Implemented machine learning pipelines for predictive analytics in healthcare
- Optimized graph-based algorithms for social network analysis applications
Senior Data Engineer
Designed and implemented robust data pipelines using PySpark and Scala for large-scale analytics. Developed and optimized Lakehouse and Medallion architectures to streamline data processing and storage. Spearheaded the migration of legacy data systems to cloud-based solutions, improving performance and scalability.
- Built real-time data ingestion frameworks with Python and SQL
- Engineered automated ETL processes for high-volume data environments
- Enhanced data governance through advanced schema management techniques
Software Engineer, Backend
Engineered backend systems using Scala with a focus on functional programming. Built fault-tolerant services with ZIO for distributed systems. Developed GraphQL APIs using Caliban for seamless client-server interactions.
- Utilized Cats library for functional programming patterns
- Optimized SQL queries for high-volume transaction systems
- Developed Python utilities for backend service monitoring