Granica Logo

Granica

Senior Software Engineer — Distributed Compute / Spark Systems

Posted 8 Days Ago
Be an Early Applicant
Hybrid
Mountain View, CA, USA
160K-240K Annually
Senior level
Hybrid
Mountain View, CA, USA
160K-240K Annually
Senior level
Build and optimize distributed compute infrastructure for large-scale analytical and AI workloads. Responsibilities include improving query execution, scheduling, resource allocation, reliability, workload routing, and compute efficiency across Spark and related systems. The role involves debugging performance bottlenecks, optimizing joins, scans, shuffles, caching, partitioning, and memory usage, and working with lakehouse formats and cloud object storage. Candidates will implement workload optimization algorithms and may contribute to open source or research.
The summary above was generated by AI
Senior Software Engineer — Distributed Compute / Spark Systems

Location: Mountain View, CA — On-site

 
About Granica

Granica builds AI infrastructure for enterprises operating massive data environments.

Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.

 

Granica’s products include:

  • Crunch — continuous optimization for enterprise lakehouse data

  • Myelin — stateful infrastructure for long-running AI agents

  • Large Tabular Models — foundation models designed for enterprise tables

 

Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

 

Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.

 
About the Role

Granica is hiring a Senior Software Engineer to build distributed compute systems for enterprise-scale data and AI workloads.

You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for distributed execution, workload optimization, query performance, scheduling, resource management, and compute cost reduction across petabyte- and exabyte-scale environments.

You will own core systems that directly affect customer compute spend, query latency, workload reliability, cluster efficiency, and the performance of large-scale analytical data processing.

This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of Spark, distributed query execution, lakehouse compute, workload scheduling, storage-aware optimization, and AI infrastructure.

You will work on distributed compute systems involving Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, Snowflake-adjacent environments, cloud object stores, and lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi.

What You’ll Do
  • Build distributed compute systems for large-scale analytical and AI workloads

  • Improve performance and cost efficiency across Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments

  • Design workload-aware systems for query execution, resource allocation, scheduling, and compute optimization

  • Optimize execution performance across joins, aggregations, scans, shuffles, spills, caching, partitioning, and task scheduling

  • Build systems that learn from workload patterns and automatically improve execution plans, cluster usage, and compute efficiency

  • Develop infrastructure for adaptive workload routing, execution planning, and data-processing reliability across large customer environments

  • Debug performance bottlenecks across query execution, metadata, storage, network, memory, CPU, and distributed compute layers

  • Work with lakehouse tables and columnar formats such as Iceberg, Delta Lake, Hudi, Parquet, and ORC to improve end-to-end workload performance

  • Build systems that reduce compute waste caused by inefficient scans, poor partitioning, small files, skew, unnecessary shuffles, and suboptimal workload placement

  • Improve reliability and failure recovery for large distributed data-processing jobs

  • Implement algorithms in workload optimization, execution efficiency, cost modeling, and data-processing performance

  • Contribute to open-source or publish research when appropriate

 
What We’re Looking For
  • Strong engineering depth in distributed systems, data processing systems, query engines, databases, or cloud infrastructure

  • Production experience with distributed compute or query systems such as Apache Spark, Spark SQL, Trino, Presto, Flink, Databricks, EMR, Glue, Hive, or similar systems

  • Hands-on experience improving performance, reliability, or cost efficiency for large-scale data-processing workloads

  • Understanding of distributed execution, query planning, scheduling, resource management, fault tolerance, and workload isolation

  • Experience with Spark internals, Spark SQL, Catalyst, Adaptive Query Execution, shuffle, joins, aggregation, spill, memory management, or task scheduling

  • Familiarity with lakehouse formats and columnar data such as Iceberg, Delta Lake, Hudi, Parquet, or ORC

  • Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of running distributed compute on top of them

  • Strong programming skills in Scala, Java, Go, Rust, C++, or similar systems-oriented languages

  • Curiosity about workload optimization, cost modeling, adaptive execution, and how compute efficiency affects AI and analytics at scale

  • A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

 
Bonus
  • Experience contributing to Apache Spark, Spark SQL, Trino, Presto, Flink, Velox, DuckDB, DataFusion, Iceberg, Delta Lake, Hudi, Parquet, ORC, or related systems

  • Experience with Catalyst, Adaptive Query Execution, cost-based optimization, query planning, vectorized execution, or distributed runtime systems

  • Experience optimizing joins, aggregations, shuffles, scans, spills, caching, partitioning, skew handling, or task scheduling

  • Experience building workload schedulers, execution control planes, resource managers, or multi-engine compute platforms

  • Experience reducing compute cost or improving workload efficiency in large-scale production data environments

  • Background in query engines, distributed runtimes, storage-aware execution, indexing, caching, encoding, compression, or adaptive query optimization

  • Research or open-source contributions in distributed systems, databases, query processing, data processing, or cloud infrastructure

 
Why Join Granica
  • Build foundational infrastructure for enterprise data and AI

  • Work on deep systems problems across distributed compute, query execution, workload optimization, scheduling, resource management, and compute efficiency

  • Partner directly with Product, Engineering, and company leadership

  • Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments

  • Work with a small, high-caliber team solving high-value infrastructure problems at massive scale

  • Have direct influence on architecture, product direction, customer outcomes, and company growth

 
Compensation & Benefits
  • Competitive salary, meaningful equity, and performance bonus for top performers

  • 401(k) with company match, comprehensive health coverage, and unlimited PTO

  • Daily catered meals in our Mountain View office

  • Support for research, publication, and conference participation

At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

 
HQ

Granica Mountain View, California, USA Office

787 Castro St, Mountain View, California, United States, 94041 2013

Similar Jobs at Granica

8 Days Ago
In-Office
Mountain View, CA, USA
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Build foundational lakehouse infrastructure for exabyte-scale AI data environments. Responsibilities include metadata and transaction systems, table maintenance, schema and partition evolution, snapshot isolation, compaction, clustering, file-layout optimization, object-store performance, columnar-format optimization, and query performance across major lakehouse engines. The role also involves debugging distributed systems, implementing compression and data-efficiency algorithms, and contributing to open-source or research efforts.
Top Skills: Amazon S3Apache FlinkApache HudiApache IcebergSparkAzure Data Lake StorageC++DatabricksDelta LakeGoGoogle Cloud StorageHive MetastoreJavaOrcParquetPrestoRustScalaSnowflakeTrinoUnity Catalog
12 Days Ago
In-Office
Mountain View, CA, USA
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Lead finance strategy including capital raising, forecasting, performance modeling, and risk management to support company’s growth and IPO readiness.
Top Skills: Ai InfrastructureFinancial Modeling
14 Days Ago
In-Office
Mountain View, CA, USA
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Develop and prototype novel diffusion and generative learning algorithms for Large Tabular Models. Design efficient training methods, representation learning techniques, benchmarks, and collaborate with academic and engineering teams to translate research into production.
Top Skills: Diffusion ModelsJaxLarge Tabular ModelsProbabilistic ModelingPythonPyTorchRepresentation LearningScalable Ml SystemsScore-Based Generative Modeling

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account