Hire
Data Systems Scala Developers
Built through years working with teams and engineers in the Scala ecosystem.
We specialize in functional backend systems, data platforms, and distributed infrastructure.
Focused, relevant introductions from a curated network.
Available Developers
Principal Agentic AI & Lead Data Systems Architect
Architected a multi-agent AI system that integrates generative AI with real-time data processing, optimizing decision-making workflows in complex environments. Spearheaded the development of a scalable graph-based reasoning engine utilizing GraphRAG and LangGraph, enhancing the efficiency of large-scale data analysis. Innovated a modular orchestration framework that leverages LangChain for seamless agent communication and coordination across distributed systems.
- Designed VertexAI-powered pipelines for real-time data ingestion.
- Implemented dynamic agent collaboration using custom orchestration protocols.
- Optimized graph traversal algorithms for high-performance data querying.
- Developed cross-platform AI models with robust fault tolerance.
Senior Software Engineer
Developed high-performance distributed systems focusing on real-time data processing and analytics, leveraging Apache DataFusion for optimized query execution. Architected scalable microservices for financial transaction processing, handling millions of operations per second with low latency. Contributed to the design of a cross-platform compiler infrastructure, enhancing multi-language interoperability and performance.
- Implemented consensus protocols for distributed databases.
- Optimized machine learning pipelines using Python and Rust.
- Led the migration of monolithic applications to cloud-native architectures.
- Enhanced JVM performance for large-scale enterprise applications.
Senior Data Engineer
- Led migration of monolithic system to microservices architecture
- Built high-throughput data pipeline processing 1M+ events per second
- Designed and implemented real-time monitoring and alerting platform
Data Engineer
Engineered large-scale data processing pipelines using Apache Spark and Scala, significantly enhancing data throughput. Developed real-time data ingestion systems with Apache Kafka, ensuring high availability and fault tolerance. Expert in streaming analytics, applying Flink for dynamic data transformations in IoT applications.
- Integrated Scala Play for responsive web data visualization
- Utilized Cats Effect for functional programming concurrency control
- Designed a custom schema registry for Kafka message validation
- Implemented end-to-end data encryption in distributed environments
Big Data Engineer
Led the design and implementation of distributed data processing systems handling petabyte-scale datasets. Spearheaded the integration of Kafka with Spark for real-time analytics pipelines in financial services. Architected a high-performance data ingestion framework using Akka to optimize processing speed and reliability.
- Developed custom SQL query optimization techniques for large datasets
- Implemented Git versioning for complex data workflows
- Pioneered Akka-based microservices for scalable data processing
- Optimized Spark jobs reducing runtime by 30%
Senior Data Engineer
Architected and managed large-scale data processing pipelines using Apache Spark, achieving real-time analytics capabilities. Led the migration of legacy systems to a cloud-based data lake architecture, improving data accessibility and scalability. Developed machine learning models for predictive analytics within the financial sector.
- Integrated Kafka streams for real-time data ingestion
- Enhanced ETL processes with Scala and functional programming
- Implemented robust data governance frameworks in Hadoop
- Optimized Spark jobs for cost-effective resource usage
Spark Performance Engineer - Data Engineer
Designed and optimized large-scale distributed data processing systems using Apache Spark, focusing on performance enhancements for data-intensive applications. Developed custom Spark transformations and actions, significantly reducing execution time for ETL pipelines in the financial analytics domain. Integrated Rust components with Scala to improve computational efficiency in real-time data processing.
- Implemented advanced caching strategies for Spark job optimization.
- Contributed to open-source Apache DataFusion project with Rust.
- Developed PySpark solutions for streaming data pipelines.
- Led migration of legacy data systems to cloud-based infrastructure.
Freelance Software | Data | ML Engineer
Developed machine learning pipelines for real-time data analysis and prediction. Deployed distributed systems on AWS Fargate for scalable processing.
- Optimized text processing with TF-IDF and k-NN
- Automated data migration using AWS DMS
- Implemented dimensionality reduction with PCA for feature optimization
Senior Scala Developer
Architected distributed systems with Scala and Kafka, achieving high throughput and fault tolerance. Built containerized applications orchestrated with Kubernetes, enhancing deployment efficiency across cloud environments. Designed infrastructure as code using Terraform, streamlining resource provisioning and management.
- Implemented event-driven architectures with ZIO.
- Migrated legacy systems to cloud-native microservices.
- Developed custom middleware solutions for data streaming pipelines.
Staff Data Engineer
Architected large-scale data pipelines using Apache Spark and Scala, optimizing for real-time analytics in the e-commerce domain. Spearheaded the migration of data infrastructure to Databricks, enhancing processing speed and reliability. Developed a custom data validation framework in Python, ensuring data integrity across distributed systems.
- Implemented Rust-based data transformation tools for performance gains
- Designed SQL-based ETL workflows for complex data models
- Automated data quality checks using custom-built Python scripts
Data Engineer
Engineered data pipelines for processing terabytes of e-commerce transaction data, enhancing data accessibility and insights. Built ETL processes for a healthcare analytics platform, improving data quality and processing efficiency. Developed a recommendation engine for a retail platform using collaborative filtering techniques.
- Automated data validation processes using Scala and Apache Spark
- Implemented data warehousing solutions with optimized query performance
- Created a scalable data ingestion framework using Kafka and Flink
- Streamlined data visualization dashboards with real-time analytics
Software Engineer / Data Engineer
Developed large-scale distributed data processing systems for real-time analytics. Specialized in building high-performance ETL pipelines and data warehousing solutions. Designed and optimized custom consensus protocols for fault-tolerant systems.
- Implemented complex event-driven architectures for financial services.
- Architected scalable microservices using Go and Rust.
- Enhanced database internals for improved query performance.
JVM Developer / Data Engineer
Designed and optimized a high-throughput real-time data processing pipeline using Kafka and Spark, achieving sub-second latency. Developed microservices architecture for a large-scale distributed system leveraging Akka and Spring frameworks.
- Implemented event-driven architectures with Apache Kafka
- Built scalable data pipelines with Apache Spark
- Optimized distributed system performance for large-scale deployments
Founding Database Engineer
Pioneered database internals and distributed systems design. Advanced knowledge in operations and systems design using AWS.
- Pioneered database internals
- Expert in distributed systems design
- Advanced AWS operations knowledge
Software Engineer, Backend
Engineered backend systems using Scala with a focus on functional programming. Built fault-tolerant services with ZIO for distributed systems. Developed GraphQL APIs using Caliban for seamless client-server interactions.
- Utilized Cats library for functional programming patterns
- Optimized SQL queries for high-volume transaction systems
- Developed Python utilities for backend service monitoring