Cogniify Logo

Cogniify

Data engineer - USA

Posted Yesterday
Remote
Hiring Remotely in United States
134K-140K Annually
Mid level
Remote
Hiring Remotely in United States
134K-140K Annually
Mid level
Build and maintain production ETL/ELT pipelines, data models, and curated datasets for analytics and AI applications. Ingest data from databases, APIs, SaaS tools, and event streams; implement quality checks, monitoring, documentation, governance, and access controls. Collaborate with analysts, data scientists, and ML engineers on reporting, feature datasets, AI search, and RAG applications. Optimize query performance and pipeline costs while using Git, code reviews, and CI/CD for safe releases.
The summary above was generated by AI

The Role

Cogniify is hiring a Mid Level Data Engineer to build reliable data pipelines and analytics datasets for reporting, business decisions and AI initiatives. You will own data work from ingestion through transformation and delivery, working with analysts, data scientists and engineers to make large datasets accurate, accessible and ready for use. This is a hands-on role for someone who enjoys both data engineering and the analytical questions behind the data.

What You Will Do

  • Build and maintain ETL/ELT pipelines using SQL, Python, dbt and tools such as Apache Spark, PySpark or Airflow.

  • Ingest data from databases, APIs, SaaS tools and event streams using connectors or custom pipelines.

  • Develop tested data models and curated datasets in Snowflake, Databricks, BigQuery or Redshift for reporting and self-service analytics.

  • Work with data scientists and ML engineers to prepare feature datasets for model training and inference.

  • Prepare and refresh structured business data that can support AI search, retrieval-augmented generation (RAG) or other Generative AI applications.

  • Build clear dashboards and analyses in Tableau, Looker, Power BI or similar tools when the work calls for it.

  • Add data quality checks, monitoring and documentation so teams can trust the data and identify pipeline issues early.

  • Improve query speed and pipeline cost; use Git, code reviews and CI/CD to release changes safely.

  • Help manage data access, lineage and sensitive information, including personally identifiable information (PII).

What We Are Looking For

  • 3 to 6 years of professional experience in data engineering, analytics engineering or a related data role with production delivery.

  • Strong SQL skills and experience writing complex transformations and improving query performance.

  • Hands-on experience with a cloud data platform. Snowflake is preferred; Databricks, BigQuery or Redshift experience is also relevant.

  • Production experience with dbt for transformation, testing and documentation.

  • Working knowledge of Python and either Pandas or PySpark for data processing.

  • Experience scheduling pipelines with Airflow, Dagster, Prefect or a similar orchestration tool.

  • A good understanding of data modeling and how to build datasets that analysts and business teams can use.

Preferred Experience

  • Apache Spark, Databricks and large-scale data processing.

  • Data quality or observability tools such as Great Expectations, Soda or Monte Carlo.

  • Streaming data with Kafka or Kinesis, or ingestion tools such as Fivetran or Airbyte.

  • Experience preparing data for ML features, AI search, embeddings or RAG applications.

  • Cloud services across AWS, Azure or Google Cloud, and data governance tools such as Unity Catalog or DataHub.

Why Join Cogniify

You will work on data products used for analytics and emerging AI applications, with room to own your pipelines and improve how teams use data. We would like to hear from engineers who care about clean data, dependable systems and useful outcomes.

Perks And Benefits Of Working With Us

  • Unlimited PTO.

  • Please ask us about our very generous parental leave, much above industry standards!.

  • Entrepreneurial culture where pushing limits and taking risks is everyday business.

  • Open communication with management and company leadership.

  • Small, dynamic teams = massive impact.

  • Medical, Dental and Vision coverage for employees.

  • Access to Disability & Life insurance.

  • Mental health and wellbeing support

  • Annual bonus program

  • Employer Stock Purchase Program (ESPP)

  • Yearly Team building experiences

  • Mentorship and sponsorship opportunities

  • Manager resources and support

We are an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other protected characteristic.

Similar Jobs

5 Days Ago
Remote or Hybrid
Sunnyvale, CA, USA
110K-286K Annually
Mid level
110K-286K Annually
Mid level
Big Data • Cloud • Logistics • Machine Learning • Retail
Designs and optimizes scalable data pipelines and architectures supporting analytics and operational needs. The role develops ETL, streaming, data quality, governance, and cloud-based workflows using modern big data technologies. Responsibilities include translating business requirements into technical solutions, monitoring data systems, improving reliability and performance, and providing technical guidance on data modeling, storage, and metadata management.
Top Skills: Apache HiveApache KafkaSparkCi/CdData ModelingData WarehousingETLGCPMetadata ManagementPython
7 Days Ago
In-Office or Remote
99K-168K Annually
Senior level
99K-168K Annually
Senior level
Consulting
Develops Scala and Python data solutions for Medicare and Medicaid claims cost measures. Responsibilities include ETL, data analysis, SQL querying, validation, quality checks, documentation, requirements communication, and collaboration with external partners. The role participates in Scrum ceremonies, test meetings, and Agile planning while supporting CMS’s Quality Payment Program. It is remote within the United States, follows East Coast hours, and may require occasional travel for planning events.
Top Skills: SparkConfluenceGitGitJIRAPysparkPythonScalaScaled Agile FrameworkScrumSparkrSQL
9 Days Ago
Remote
USA
135K-150K Annually
Senior level
135K-150K Annually
Senior level
Healthtech
Build and operate Danaher’s enterprise data platform automation and reliability layer. Responsibilities include Infrastructure-as-Code, CI/CD, self-service developer enablement, data platform provisioning, observability, SLOs, auto-remediation, FinOps, compliance guardrails, and event-driven operations across Snowflake, Azure, Matillion, dbt, and Airflow. The role also develops AI-driven operators and copilots using Azure AI Foundry, Anthropic Claude, and related agent frameworks.
Top Skills: Anthropic ClaudeApache AirflowAutogenAzureAzure Ai FoundryAzure MonitorAzure OpenaiDbtGithub ActionsGitlabLanggraphLog AnalyticsMatillionOpenaiPrompt FlowPulumiPythonSemantic KernelServicenowSnowflakeSnowflake CortexTerraform

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account