Tether.io Logo

Tether.io

AI Inference Engineer QVAC (100% remote Worldwide)

Reposted 8 Days Ago
In-Office or Remote
Hiring Remotely in Dubai
Mid level
In-Office or Remote
Hiring Remotely in Dubai
Mid level
As an AI Inference Engineer, you will optimize C++ systems for AI model inference on edge devices, collaborating with researchers to deploy and enhance models while ensuring runtime stability and performance.
The summary above was generated by AI

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the role:

You will own the inference backbone behind QVAC's local AI stack: the C++ systems layer that makes models run fast, reliably, and predictably on real user hardware. The role is centered on engineering quality at runtime level, including startup behavior, memory pressure, throughput/latency balance, and long-session stability. You will define and evolve the core abstractions that inference features depend on, so new capabilities can be added without sacrificing performance or maintainability. This is a role for someone who enjoys low-level problem solving, clear technical ownership, and building infrastructure that other teams trust in production. Your work directly enables private, on-device AI experiences and helps set the technical foundation for QVAC's next generation of peer-to-peer AI products.

About the job

You'll work on the C++ layer that powers local AI, porting and enhancing inference engines like llama.cpp or similar, to run efficiently on edge devices. Your focus is on the runtime: making models load faster, run leaner, and perform well across different hardware. You'll ensure that the inference layer is stable, optimized, and ready for integration with the rest of the stack.

This role is for engineers who want to work close to the metal, enabling private and fast on-device AI without relying on cloud infrastructure.

Responsibilities

  • Work on deploying machine learning models to edge devices using the frameworks: llama.cpp, ggml

  • Collaborate closely with researchers to assist in coding, training and transitioning models from research to production environments

  • Integrate AI features into existing products, enriching them with the latest advancements in machine learning

  • Strong programming skills in C++

  • Strong experience with Llama.cpp and ggml inference engines, which facilitates the deployment of models to specific GPU architectures

  • Experience with any GPU framework between Cuda, Vulkan, Metal, OpenCL

  • Good understanding of deep learning concepts and model architectures

  • Experience with transformers, LLMs, Diffusion models

  • Demonstrated ability to rapidly assimilate new technologies and techniques

  • A degree in Computer Science, AI, Machine Learning, or a related field, complemented by a solid track record in AI R&D

Bonus points if:

  • You know how to train/fine-tune a LLM

  • You have productionized models

  • You have research experience in new model architectures

  • You have experience with distributed systems

  • Javascript experience

Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @tether.to or @tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Similar Jobs

2 Hours Ago
In-Office or Remote
Expert/Leader
Expert/Leader
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Lead market intelligence, stakeholder engagement, and ecosystem growth for Circle across North Africa. Drive compliant market expansion, coordinate cross-functional execution, support regulatory outreach, identify high-potential use cases and partners, and represent Circle with senior regional stakeholders.
Top Skills: Apple MacosChatgptGeminiGoogle SuiteSlack
2 Hours Ago
Remote
Senior level
Senior level
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Serve as the L2 SME for Netcracker Convergent Billing & Rating, handling P1/P2 incidents, RCA, troubleshooting billing/rating/charging issues, validating data consistency across catalog and engines, coordinating escalations and releases, maintaining runbooks, and mentoring L1 teams to meet SLA/OLA targets.
Top Skills: Billing EngineCdr IngestionEsbJavaJavaScriptJIRALinuxMediation PlatformsNetcrackerNetcracker Convergent Billing & RatingProduct CatalogProduct Offering (Po)Product Structure (Ps)Rating EngineRemedyRest ApisRule EngineServicenowShell ScriptingSoap ApisSQLSsl CertificatesUnix
15 Hours Ago
Remote or Hybrid
Mid level
Mid level
Artificial Intelligence • Information Technology • Software • Analytics • Consulting • Generative AI
The Forward Deployed Software Engineer will build applications for clients, leveraging customer data to enhance business operations and deliver impactful results. Responsibilities include requirement gathering, independent problem-solving, and collaboration with teams on implementation and outcomes.
Top Skills: AipPalantir FoundryPythonReactSparkTypescript

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account