Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Engineer Staff Data Engineer at Code for America

Leads technical execution of data infrastructure and pipelines for government services, ensuring data reliability and usability at scale across multiple systems.

Lead Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

Code for America believes government can work for the people, by the people, in the new digital age, and that government at all levels can and should work well for all people. For more than a decade, we’ve worked to show that with the mindful use of technology, we can break down barriers, meet community needs, and find real solutions.

Our employees build and transform government and community tools and services, making them so good they inspire change. We merge the best parts of technology, nonprofit, and government to help support the people who need it most.

With a focus on transparency and fairness, and deep empathy for partners in government and community organizations and the people that our partners serve, we’re building a movement of motivated change agents driven by meaningful results and lasting impact.

At Code for America, you contribute to exciting work while learning and developing in a supportive and flexible environment. Our compensation and benefits are holistic and thoughtfully curated to represent our employees and our mission. Help us drive real generational change that lasts.

Code for America is looking for a talented Staff Data Engineer who will lead the technical execution of our data infrastructure behind high-impact government services. In this role you will make data reliable and usable at scale — improving how our teams collect, move, model, and analyze the data that tells us whether the services we build actually work for the people who need them most.

The pace of change in the policy and government landscape is accelerating at the same time that technology possibilities are shifting rapidly. Together, these shifts create new windows of opportunity to develop technical solutions that make critical public services more effective for the people who need them. Staff data engineers keep delivery moving through those moments, supporting staff and holding high delivery standards so that data are trustworthy, and reachable by the people who need it.

About the role:

Code for America’s Staff Data Engineers build and operate the infrastructure that moves data through the products we ship, between our systems and our partners’ systems, and into the analysis that drives meaningful improvements in outcomes across the tax, criminal justice, and safety net delivery areas. Staff Data Engineers lead the technical execution of data infrastructure projects and solutions while guiding their teams’ technical direction, setting a high standard for technical judgment, and enabling other data scientists and engineers through collaboration and mentorship.

This role will report to a Data Science Manager and is expected to travel no more than 10 % of the time.

Code for America is based in California and can employ those who reside full-time within the United States. This is a remote position.

In this position you will:

  • Data Infrastructure & Pipeline Engineering

    • Lead cross-functional technical planning with Data Scientists, Software Engineers, and Product Managers to translate data architecture requirements into scalable technical solutions for your team’s projects
    • Create new data pipelines, data transfers, and compliance-oriented infrastructure to facilitate seamless data utilization within cloud environments
    • Contribute to best practices for dataset management and optimization, ensuring data consistency and performance
    • Implement data transformation, integration, and validation processes to support analytics and reporting needs
    • Partner with our DevOps team to apply deployment standards, monitoring frameworks, and operational best practices for data infrastructure
    • Lead technical investigations into data issues within the systems you own, resolving them with thorough communication and recommending process changes where necessary.
    • Optimize and fine-tune data pipelines for improved speed, reliability, and efficiency
    • Lead generation and optimization of data architecture recommendations and demonstrate the ability to implement them
    • Champion privacy, security, and data governance practices throughout the entire data lifecycle
    • Support Engineering’s culture of writing by producing quality documentation and learning materials.
  • Technical Leadership & Enablement

    • Contribute to the development of data engineering best practices and standards, and reinforce them across the team.
    • Collaborate with Product Managers and other cross-functional peers; adapt your approach as priorities, technology, and team needs shift; and actively help others do the same when facing major changes in direction or tooling.
    • Mentor engineers and data scientists on data engineering best practices, pipeline design, and sound architectural thinking, and create conditions for them to grow into unfamiliar problem spaces through pairing, documentation, or direct coaching.
  • Other duties as assigned

    • Engage in ad hoc projects and initiatives as required by the organization’s evolving needs.

About you:

  • 7+ years of experience in data engineering or related technical fields
  • Expert-level SQL and experience with python, R
  • Experience with Postgres and cloud-based data warehousing (e.g. BigQuery, Snowflake)
  • Expert at data orchestration tools (e.g. Dagster) for building and managing data pipelines
  • Expert at using dbt for data transformation and modeling within data warehouses
  • Demonstrated experience with Metabase or similar business intelligence tools (e.g. Tableau, PowerBI, Looker)
  • Experience in ensuring that data pipelines are accessible to Data Scientists through notebook tools (e.g. Hex)
  • Experience with cloud infrastructure and infrastructure as code.
  • Excellent collaboration skills, with a demonstrated ability to manage cross-functional relationships, respond to stakeholder needs, and facilitate effective decision-making conversations
  • Excellent written and verbal communication skills in English.
  • Demonstrated ability to use AI coding tools judiciously, validating and refining generated code for quality, and modeling that discipline for teammates is needed for this specific role; organizationally, a curiosity about emerging AI tools and a commitment to using them responsibly and effectively
  • A deep commitment to Code for America’s mission of making government services work well for the people who need them most.

It’s a bonus if you have:

  • Experience with making government services better for people who need them most
  • Personal or professional experience with the social safety net or other mission-relevant government services.
  • Experience building and integrating AI-powered features in production systems, with evaluation and monitoring.

What you’ll get -

Salary:

Code for America’s salary bands are transparent as a part of our commitment to transparency and fairness. As part of our hiring practices, we aim to target the midpoint of the 2nd quartile of the range for all new hires.

Offer targets vary based on market / geographic location. The offer targets for this role range from $128,945 to $157,850,annually. Your Recruiter will discuss compensation in more detail & answer any related questions, during the initial Recruiter phone screen.

Benefits and perks:

This role includes a comprehensive benefits package including the following:

  • Values:
    • Leadership and teammates who share a strong work ethic and values, and who respect and care for one another
    • A collaborative, cross-functional, hardworking, and joyful environment
  • Employee Enablement Support:
    • Laptop provided
    • A one-time $700 payment for remote environment setup; $200 stipend (in first paycheck) and up to $500 reimbursement, in accordance with our equipment policy
    • Cell phone and/or internet reimbursement of $50 per month
  • Professional Development:
    • $500 annual (per calendar year) stipend towards professional development; prorated at time of hire
    • Up to $500 of professional development funds can be rolled over each year, up to a maximum of $1000
    • Training / guidance for staff required to utilize AI as part of their role, plus opportunities for employees to gain AI-related skills to support job and career growth
  • Retirement & 401k Plans:
    • Employees receive a 100% employer match on the first 3% of contributions.
    • Employees with 3+ years of service receive an additional 50% match on contributions between 3% and 5%, for a maximum employer contribution of 5%
  • Medical:
    • At least one no cost health insurance option for full-time employees for employee-only coverage
    • A minimum of 80% of the cost of dependent coverage
  • Remote Work:
    • Code for America employees may work remotely across the US
    • Code for America employees main residence must be within the US
    • Full-time employees work 40 hours per week, Monday - Friday
    • Collaborative working hours: we aim to hold all internal meetings between 10 AM - 3 PM PT. We expect all Code for America staff to be available during these set working hours
  • Time Off:
    • Open personal time off ( subject to manager approval), a minimum of 14 paid holidays, and an org-wide closure from Christmas Day through New Year’s Day
    • Paid sick time; up to 96 hours annually
    • 17 weeks of paid parental and family leave
    • 3 weeks of paid sabbatical after 5 years of service

Timeline:

From the time you are contacted by a Recruiter, you can expect a 1-2 month, or longer, hiring process. The specific hiring process and timeline goals for this role will be covered during the initial Recruiter phone screen. Code for America’s standard interview process details are also available on our careers page.

Equal Employment Opportunity:

Code for America is an equal opportunity employer. Applicants will not be discriminated against because of race, color, creed, sex, sexual orientation, gender identity or expression, age, religion, national origin, citizenship status, disability, ancestry, marital status, veteran status, medical condition or any protected category prohibited by local, state or federal laws.

Code for America Workers United:

This position is covered by a Collective Bargaining Agreement between Code for America and Code for America Workers United, affiliated with OPEIU, Local 1010. The agreement was ratified on January 13, 2026, and is currently in effect.

#LI-MD1

#LI-Remote

Read the full description
Engineer Senior/Staff Machine Learning Engineer (Model Dev) at Artera.net

Lead ML engineer develops and validates AI biomarkers for cancer care, owning models from conception through regulatory submission and production deployment.

Lead Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

About Us: Artera is an artificial intelligence company dedicated to transforming cancer care. We’ve developed foundation models that analyze clinical and pathology data, generating actionable insights that guide therapy selection and improve outcomes for cancer patients. By continuously improving these models, we aim to uncover the biological mechanisms driving cancer progression.

We’re looking for an experienced machine learning engineer to own AI biomarker development end to end — from problem framing with clinical and biostatistics partners, through model development and validation, to regulatory submission and production deployment. Beyond owning a biomarker program, you’ll take on the hardest cross-cutting problems in our field: robustness across scanners and sites, mechanistic interpretability of model decisions, and the next generation of our pathology foundation models.

Essential Responsibilities:

  • Lead the technical effort and define the strategic vision for patient-facing products, in partnership with product, biostatistics, clinical development, and regulatory/quality.

  • Design and build AI-based biomarkers on multimodal data — including whole-slide images, clinical variables, and molecular data — to predict patient outcomes, treatment benefit, and molecular traits.

  • Advance our core self-supervised foundation models and the downstream architectures built on them (multiple-instance learning, time-to-event / hazard models, segmentation and classification components), with generalization as a first-order objective.

  • Own score reproducibility across scanners, institutions, staining protocols, and patient populations.

  • Develop and integrate mechanistic interpretability methods to explain model decisions, build clinician trust, and drive actionable model improvements.

  • Architect tools and processes that streamline the end-to-end model development lifecycle — from prototyping through production deployment and monitoring — ensuring efficiency, reproducibility, regulatory compliance, and scale.

  • Author and defend regulatory and quality documentation, and represent AI in design and development reviews.

  • Plan and manage delivery: break multi-quarter programs into milestones, manage dependencies across AI, platform, biostatistics, and clinical teams, surface risk early, and hold submission and launch dates.

  • Publish in peer-reviewed journals and present at clinical and ML venues; support external academic and industry collaborations.

  • Mentor and coach machine-learning scientists and engineers, fostering their technical growth and collaboration skills, and raise the bar on scientific rigor, code quality, and written communication across the team.

Experience Requirements:

  • 5+ years of industry experience building deep learning systems in PyTorch (or TensorFlow).

  • 2+ years of experience as a technical lead, launching and monitoring machine-learning products in production environments.

  • Demonstrated depth in oncology and biomarker development: familiarity with cancer biology and treatment pathways, clinical endpoints, risk stratification, and what makes a biomarker clinically actionable.

  • Demonstrated project management ability — scoping, sequencing, and managing dependencies and risk across multiple teams on dated deliverables.

  • Proven ability to communicate complex ML concepts effectively to cross-functional, non-ML collaborators.

  • Experience mentoring or managing ML scientists and engineers.

Desired:

  • Experience building ML on complex clinical data — medical imaging, multi-omics, or longitudinal patient records — including weakly supervised learning and handling variation across sites, devices, and protocols.

  • Experience developing ML in a regulated environment — FDA 510(k)/De Novo, CE/UKCA, SaMD, design controls, or CLIA/LDT validation.

  • Experience with self-supervised representation learning (e.g., DINOv3) and adapting medical foundation models to downstream clinical tasks.

  • Experience with data from randomized controlled trials and multi-institutional clinical cohorts.

  • Peer-reviewed publications and conference presentations; history of external academic or industry collaborations.

  • Experience with cloud-scale training and workflow orchestration (e.g., Flyte / Union, Kubernetes, AWS), experiment tracking, and reproducible ML pipelines.

$180,000 - $240,000 a year

In addition to base salary, equity is a core component of our compensation. We also offer 401k matching, unlimited paid time off (PTO), and more.

The base salary is competitive and commensurate with experience, qualifications, and other factors to be discussed during the interview process.

Equal Employee Opportunity:At Artera, we value bringing together individuals from diverse backgrounds to develop new andinnovative solutions for patients and physicians. As an equal opportunity employer, we do notdiscriminate on the basis of race, color, religion, national origin, age, sex (including pregnancy),physical or mental disability, medical condition, genetic information gender identity orexpression, sexual orientation, marital status, protected veteran status, or any other legallyprotected characteristic.

Read the full description
Engineer Tech Lead - Code Plane [IC5] at Sourcegraph

Lead technical architecture and product development for code infrastructure platform serving AI agents and engineering teams at scale.

Lead Remote Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

Who we are

Our mission is to bring clarity and control to the world’s most complex codebases. AI is accelerating code creation, but the infrastructure to understand, oversee, and evolve that code hasn’t kept pace. Sourcegraph gives engineering organizations full visibility across their systems, precise context for their agents, and the ability to execute coordinated code changes at scale. As agentic development becomes the dominant engineering paradigm, we provide the context layer teams need to take control of their codebase.

With Code Search, Deep Search, MCP, and Agentic Batch Changes, we deliver on that mission today - giving engineering teams and their AI tools the cross-repo context to navigate massive codebases with confidence, and the ability to make changes across hundreds of repositories at once.

Companies like Stripe, Reddit, and Leidos rely on Sourcegraph to ship faster and with higher quality. We’re backed by a16z, Sequoia, and Redpoint, and proud to operate as a globally distributed team that values high agency, direct communication, and customer love.

If you want to build the infrastructure that lets every engineering team - and every agent they deploy - operate on their codebase with confidence, join us.

Hours & location

🌎 While we hire almost anywhere in the world, we have a preference for someone to reside in the following locations for this role. However, if you feel qualified, we welcome you to apply regardless of location. No matter what, working hours must overlap with CEST for at least 20 hours/week.

Preferred locations:

  • Europe
  • EST

Why this job is exciting

The Code Plane team owns the Sourcegraph products that help developers - and the agents working on their behalf - take action on code, at scale, across the world’s largest codebases. Think “data plane” and “control plane” for enterprise code: the surfaces where intent becomes code changes across hundreds (or more!) of repositories.

This is a team that builds products that use AI and products for AI. The roadmap, the architecture, and the day-to-day technical decisions all hinge on a solid experience with and a clear-eyed view of what AI models and agents are good at, where they fail, and how to design products around them. A strong handle on and opinion of AI, formed from real-world, hands-on use, not just observation, is a hard requirement for this role. If you’re excited to ship for the AI agent ecosystem rather than simply watch from the sidelines, you’ll find a lot to love here.

We’re hiring you to be the technical leader around whom this team rallies. Our product surfaces are evolving drastically as agents reshape how software is built, and we are adjusting our engineering teams to match. It is an exciting time, and a rare chance to be a strong leader shaping dev tools for the agentic age of coding. Code Plane owns high-stakes, fast-moving products at the center of that shift and needs a tech lead who will set technical and product direction, drive the roadmap from issue to shipped, and keep the team unblocked and moving. You are a generalist by choice. You go where the problem is, backend, AI agent, or frontend, and you are who people come to when it crosses a boundary.

The team’s surface areas include:

  • Batch Changes - our battle-tested execution layer for massive cross-repo changes helps developers automate large-scale code changes across all of an enterprise’s repositories: keeping code up-to-date, fixing critical security issues, and paying down tech debt, with every change tracked through checks and code review until it merges. Teams at Stripe, Uber, and Dropbox, and others use it to drive migrations, refactors, and security fixes.
  • Agentic Batch Changes. The frontier agent for code changes at scale. Building on the foundation of Batch Changes, an outer-loop agent writes  code-modification programs and orchestrates coding agents to roll out changes across the world’s largest codebases, complete with CI feedback loops and post-publish remediation. Anything a single coding agent can do in one repository, Agentic Batch Changes drives across all of them at once, from thousands of repositories to the largest monorepos.
  • Code Monitors and Code Insights - our “signals over code” products. Code Monitors notify and alert based on detected changes in critical code. A leaked secret, a deprecated library, an anti-pattern creeping back in can all turn into alerts. Code Insights turns a codebase into data to back up business decisions and customizable dashboards surfacing critical info about repos so devs can track what matters over time - migration progress, dependency adoption, vulnerability mitigation, code health - with data-driven answers.
  • The src CLI tool - the command-line surface that humans and the agents they employ use to automate Sourcegraph in their workflows. With a natural, scriptable surface, it’s quickly becoming the connective tissue between Sourcegraph and the agent ecosystem.

The features your team ships are used directly by developers and by agents working on their behalf, and you’ll partner closely with Product, Design, Customer Engineering, and adjacent engineering teams to turn customer pain into shipped product.

📅 Within one month, you will…

  • Build relationships across your team and familiarize yourself with its surface area - Batch Changes, Agentic Batch Changes, Code Monitors, Code Insights, src CLI - by running them as a user and an agent would, and form a point of view on the most important product and technical bets to make next.
  • Meet your cross-functional partners in Product, Design, Solutions Engineering, and adjacent engineering teams, and start translating customer feedback into a concrete view of what the team should ship.
  • Join the team’s on-call support rotation.

📅 Within three months, you will…

  • Own a meaningful slice of Agentic Batch Changes end-to-end and ship critical pieces yourself.
  • Be the technical lead for the team’s roadmap, scoping and sequencing projects across the team’s product surfaces, setting clear expectations for ownership, and clarifying priorities.
  • Strengthen the team practices that make their work reliable: how we design, review, test, and deliver.
  • Mentor teammates and continue to build high-trust relationships.

📅 Within six months , you will…

  • Be the recognized technical visionary for Code Plane: to your teammates, and increasingly the wider department, defer to on architecture, quality, and product trade-offs in this space.
  • Be known across the company as the person who deeply understands what customers and agent ecosystems need from Code Plane, and where Sourcegraph should invest next.
  • Be setting the direction of the team’s roadmap with conviction backed by evidence, and mobilizing other engineers to execute on it.

About you

You are a technical leader and software engineer whom people want to follow. You set technical direction, make the hard architecture and tradeoff calls, and keep the team unblocked and moving. You think in terms of customers and outcomes, not tickets. You bring a technical perspective to what the team should and should not take on, and you make the quality bar real: review culture, testing standards, and release practices. You’re scoping with Product and Design and making calls on what to cut so the team can ship. You’re tackling our hardest technical problems hands-on, and also guiding your team to grow.

Agents are one of the most important parts of this job. You have shipped for the agent ecosystem: something that agents or other programs call, with evals behind it, so you know whether a prompt, tool, or model change made things better or worse. You can say precisely where agents fail, because you have encountered it yourself and measured it.

You can take a position, argue it persuasively, and bring people with you, including the ones who started out disagreeing, and you say plainly what would change your mind. You contribute to a collaborative, respectful, async-first culture, and you keep stakeholders informed without being asked.

  • 8+ years of professional software engineering experience
  • Fluency across our stack, backend and frontend: Go, TypeScript, SvelteKit and React, GraphQL APIs, PostgreSQL, and Docker and Kubernetes in a multi-service environment.
  • A strong technical background,  with the ability to provide technical guidance on architecture, code quality, and trade-offs.
  • Proven experience on teams that build end-user product features — ideally developer tools, agent/AI products, CLIs, or other technical product surfaces.
  • Experience partnering closely with Product and Design to take features from idea to shipped, and incorporating customer feedback into the roadmap.
  • You have built something that agents or other programs call, not just something you used an agent to build. MCP servers, agent SDKs, tool surfaces, harnesses, or a CLI designed for non-human callers.
  • Hands-on experience using and reasoning about coding agents (Amp, Claude Code, Cursor, or similar), with strong opinions about what good looks like when humans and agents share a workflow.
  • Async-first communication skills and experience working in a globally distributed remote team.

Strongly preferred

  • You have orchestrated autonomous work at scale and built the machinery that survives it: failures that resume rather than restart, costs that stay inside a budget, and a blast radius you chose deliberately rather than discovered.
  • Experience owning a service you did not build, with paying customers on a schema you inherited, and making it better without a rewrite.
  • A track record of raising the bar for other engineers, such as a testing or code review standard or working with AI assistants
  • You are comfortable making architectural decisions where no established pattern exists, and equally comfortable revising them when the evidence turns, without needing to have been right the first time.
  • You can decide what good looks like for a system that has no single correct answer, build the measurement that proves it, and hold that standard as the ground shifts underneath you.
  • You can sit with a customer, listen past the symptom, and come back with a technical bet the team can build.
  • Experience as a Team or Tech Lead on a software engineering team.

Level

📊 This job is an IC5.  You can read more about our job leveling philosophy in our Handbook.

Compensation

💸 We pay above-market salaries because we want to hire exceptional people who can focus on building great products, not worrying about paying bills. As an open and transparent company, our compensation philosophy and pay bands are visible to every Sourcegraph teammate, and we strive to make our approach equitable, explainable, and competitive.

Your base salary is determined by the IC5 pay band for your location zone (1-4). Our pay bands are informed by market data and designed to ensure competitive compensation wherever you live. During the recruiting process, we’ll discuss the range applicable to you based on job level, relevant skills, experience, qualifications, and location zone.

💰 The starting salary for the IC5 pay band in each zone is:

  • Zone 2: $192,000 USD
  • Zone 3: $144,000 USD

📈 In addition to competitive cash compensation, we offer meaningful equity (because when Sourcegraph succeeds, we want you to succeed, too) and generous perks & benefits.

Interview process

Below is the interview process you can expect for this role (you can read more about the types of interviews in our Handbook). It may look like a lot of steps, but rest assured that we move quickly and the steps are designed to help you get the information needed to determine if we’re the right fit for you… Interviewing is a two-way street, after all!

We expect the interview process to take 4.75 hours in total.

👋 Introduction Stage - we have initial conversations to get to know you better…

  • [30m] Recruiter Screen
  • [60m] Hiring Manager Screen / Resume Deep Dive

🧑‍💻 Team Interview Stage - we then delve into your experience in more depth and introduce you to members of the team, including cross-functional partners…

  • [60m] System Design
  • [60m] Harden a code research agent
  • [60m] Cross-functional team collaboration / Values

🎉 Final Interview Stage- we move you to our final round, where you gain a better understanding of our business and values holistically…

  • [30] Leadership
  • We check references and conduct your background check

Please note - you are welcome to request additional conversations with anyone you would like to meet, but didn’t get to meet during the interview process.

Learn more about us

You can learn more about what it is like to work at Sourcegraph by reading our handbook.

We are an ambitious team who are collectively working hard to build the most influential company in the world. You can read more about our culture, competitive compensation and benefits here.

Sourcegraph is an equal opportunity workplace; we welcome people from all backgrounds.

Sourcegraph participates in E-Verify for U.S. Employees.

Read the full description
Engineer Staff Data Engineer at Code for America

Leads technical execution of data infrastructure projects, ensuring data reliability and usability across government services and systems.

Lead Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

Code for America believes government can work for the people, by the people, in the new digital age, and that government at all levels can and should work well for all people. For more than a decade, we’ve worked to show that with the mindful use of technology, we can break down barriers, meet community needs, and find real solutions.

Our employees build and transform government and community tools and services, making them so good they inspire change. We merge the best parts of technology, nonprofit, and government to help support the people who need it most.

With a focus on transparency and fairness, and deep empathy for partners in government and community organizations and the people that our partners serve, we’re building a movement of motivated change agents driven by meaningful results and lasting impact.

At Code for America, you contribute to exciting work while learning and developing in a supportive and flexible environment. Our compensation and benefits are holistic and thoughtfully curated to represent our employees and our mission. Help us drive real generational change that lasts.

Code for America is looking for a talented Staff Data Engineer who will lead the technical execution of our data infrastructure behind high-impact government services. In this role you will make data reliable and usable at scale — improving how our teams collect, move, model, and analyze the data that tells us whether the services we build actually work for the people who need them most.

The pace of change in the policy and government landscape is accelerating at the same time that technology possibilities are shifting rapidly. Together, these shifts create new windows of opportunity to develop technical solutions that make critical public services more effective for the people who need them. Staff data engineers keep delivery moving through those moments, supporting staff and holding high delivery standards so that data are trustworthy, and reachable by the people who need it.

About the role:

Code for America’s Staff Data Engineers build and operate the infrastructure that moves data through the products we ship, between our systems and our partners’ systems, and into the analysis that drives meaningful improvements in outcomes across the tax, criminal justice, and safety net delivery areas. Staff Data Engineers lead the technical execution of data infrastructure projects and solutions while guiding their teams’ technical direction, setting a high standard for technical judgment, and enabling other data scientists and engineers through collaboration and mentorship.

This role will report to a Data Science Manager and is expected to travel no more than 10 % of the time.

Code for America is based in California and can employ those who reside full-time within the United States. This is a remote position.

In this position you will:

  • Data Infrastructure & Pipeline Engineering

    • Lead cross-functional technical planning with Data Scientists, Software Engineers, and Product Managers to translate data architecture requirements into scalable technical solutions for your team’s projects
    • Create new data pipelines, data transfers, and compliance-oriented infrastructure to facilitate seamless data utilization within cloud environments
    • Contribute to best practices for dataset management and optimization, ensuring data consistency and performance
    • Implement data transformation, integration, and validation processes to support analytics and reporting needs
    • Partner with our DevOps team to apply deployment standards, monitoring frameworks, and operational best practices for data infrastructure
    • Lead technical investigations into data issues within the systems you own, resolving them with thorough communication and recommending process changes where necessary.
    • Optimize and fine-tune data pipelines for improved speed, reliability, and efficiency
    • Lead generation and optimization of data architecture recommendations and demonstrate the ability to implement them
    • Champion privacy, security, and data governance practices throughout the entire data lifecycle
    • Support Engineering’s culture of writing by producing quality documentation and learning materials.
  • Technical Leadership & Enablement

    • Contribute to the development of data engineering best practices and standards, and reinforce them across the team.
    • Collaborate with Product Managers and other cross-functional peers; adapt your approach as priorities, technology, and team needs shift; and actively help others do the same when facing major changes in direction or tooling.
    • Mentor engineers and data scientists on data engineering best practices, pipeline design, and sound architectural thinking, and create conditions for them to grow into unfamiliar problem spaces through pairing, documentation, or direct coaching.
  • Other duties as assigned

    • Engage in ad hoc projects and initiatives as required by the organization’s evolving needs.

About you:

  • 7+ years of experience in data engineering or related technical fields
  • Expert-level SQL and experience with python, R
  • Experience with Postgres and cloud-based data warehousing (e.g. BigQuery, Snowflake)
  • Expert at data orchestration tools (e.g. Dagster) for building and managing data pipelines
  • Expert at using dbt for data transformation and modeling within data warehouses
  • Demonstrated experience with Metabase or similar business intelligence tools (e.g. Tableau, PowerBI, Looker)
  • Experience in ensuring that data pipelines are accessible to Data Scientists through notebook tools (e.g. Hex)
  • Experience with cloud infrastructure and infrastructure as code.
  • Excellent collaboration skills, with a demonstrated ability to manage cross-functional relationships, respond to stakeholder needs, and facilitate effective decision-making conversations
  • Excellent written and verbal communication skills in English.
  • Demonstrated ability to use AI coding tools judiciously, validating and refining generated code for quality, and modeling that discipline for teammates is needed for this specific role; organizationally, a curiosity about emerging AI tools and a commitment to using them responsibly and effectively
  • A deep commitment to Code for America’s mission of making government services work well for the people who need them most.

It’s a bonus if you have:

  • Experience with making government services better for people who need them most
  • Personal or professional experience with the social safety net or other mission-relevant government services.
  • Experience building and integrating AI-powered features in production systems, with evaluation and monitoring.

What you’ll get -

Salary:

Code for America’s salary bands are transparent as a part of our commitment to transparency and fairness. As part of our hiring practices, we aim to target the midpoint of the 2nd quartile of the range for all new hires.

Offer targets vary based on market / geographic location. The offer targets for this role range from $128,945 to $157,850,annually. Your Recruiter will discuss compensation in more detail & answer any related questions, during the initial Recruiter phone screen.

Benefits and perks:

This role includes a comprehensive benefits package including the following:

  • Values:
    • Leadership and teammates who share a strong work ethic and values, and who respect and care for one another
    • A collaborative, cross-functional, hardworking, and joyful environment
  • Employee Enablement Support:
    • Laptop provided
    • A one-time $700 payment for remote environment setup; $200 stipend (in first paycheck) and up to $500 reimbursement, in accordance with our equipment policy
    • Cell phone and/or internet reimbursement of $50 per month
  • Professional Development:
    • $500 annual (per calendar year) stipend towards professional development; prorated at time of hire
    • Up to $500 of professional development funds can be rolled over each year, up to a maximum of $1000
    • Training / guidance for staff required to utilize AI as part of their role, plus opportunities for employees to gain AI-related skills to support job and career growth
  • Retirement & 401k Plans:
    • Employees receive a 100% employer match on the first 3% of contributions.
    • Employees with 3+ years of service receive an additional 50% match on contributions between 3% and 5%, for a maximum employer contribution of 5%
  • Medical:
    • At least one no cost health insurance option for full-time employees for employee-only coverage
    • A minimum of 80% of the cost of dependent coverage
  • Remote Work:
    • Code for America employees may work remotely across the US
    • Code for America employees main residence must be within the US
    • Full-time employees work 40 hours per week, Monday - Friday
    • Collaborative working hours: we aim to hold all internal meetings between 10 AM - 3 PM PT. We expect all Code for America staff to be available during these set working hours
  • Time Off:
    • Open personal time off ( subject to manager approval), a minimum of 14 paid holidays, and an org-wide closure from Christmas Day through New Year’s Day
    • Paid sick time; up to 96 hours annually
    • 17 weeks of paid parental and family leave
    • 3 weeks of paid sabbatical after 5 years of service

Timeline:

From the time you are contacted by a Recruiter, you can expect a 1-2 month, or longer, hiring process. The specific hiring process and timeline goals for this role will be covered during the initial Recruiter phone screen. Code for America’s standard interview process details are also available on our careers page.

Equal Employment Opportunity:

Code for America is an equal opportunity employer. Applicants will not be discriminated against because of race, color, creed, sex, sexual orientation, gender identity or expression, age, religion, national origin, citizenship status, disability, ancestry, marital status, veteran status, medical condition or any protected category prohibited by local, state or federal laws.

Code for America Workers United:

This position is covered by a Collective Bargaining Agreement between Code for America and Code for America Workers United, affiliated with OPEIU, Local 1010. The agreement was ratified on January 13, 2026, and is currently in effect.

#LI-MD1

#LI-Remote

Read the full description
Engineer Tech Lead - Code Plane [IC5] at Sourcegraph

Tech lead architecting code infrastructure products that enable developers and AI agents to navigate and modify large codebases at scale.

Lead Remote Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

Who we are

Our mission is to bring clarity and control to the world’s most complex codebases. AI is accelerating code creation, but the infrastructure to understand, oversee, and evolve that code hasn’t kept pace. Sourcegraph gives engineering organizations full visibility across their systems, precise context for their agents, and the ability to execute coordinated code changes at scale. As agentic development becomes the dominant engineering paradigm, we provide the context layer teams need to take control of their codebase.

With Code Search, Deep Search, MCP, and Agentic Batch Changes, we deliver on that mission today - giving engineering teams and their AI tools the cross-repo context to navigate massive codebases with confidence, and the ability to make changes across hundreds of repositories at once.

Companies like Stripe, Reddit, and Leidos rely on Sourcegraph to ship faster and with higher quality. We’re backed by a16z, Sequoia, and Redpoint, and proud to operate as a globally distributed team that values high agency, direct communication, and customer love.

If you want to build the infrastructure that lets every engineering team - and every agent they deploy - operate on their codebase with confidence, join us.

Hours & location

🌎 While we hire almost anywhere in the world, we have a preference for someone to reside in the following locations for this role. However, if you feel qualified, we welcome you to apply regardless of location. No matter what, working hours must overlap with CEST for at least 20 hours/week.

Preferred locations:

  • Europe
  • EST

Why this job is exciting

The Code Plane team owns the Sourcegraph products that help developers - and the agents working on their behalf - take action on code, at scale, across the world’s largest codebases. Think “data plane” and “control plane” for enterprise code: the surfaces where intent becomes code changes across hundreds (or more!) of repositories.

This is a team that builds products that use AI and products for AI. The roadmap, the architecture, and the day-to-day technical decisions all hinge on a solid experience with and a clear-eyed view of what AI models and agents are good at, where they fail, and how to design products around them. A strong handle on and opinion of AI, formed from real-world, hands-on use, not just observation, is a hard requirement for this role. If you’re excited to ship for the AI agent ecosystem rather than simply watch from the sidelines, you’ll find a lot to love here.

We’re hiring you to be the technical leader around whom this team rallies. Our product surfaces are evolving drastically as agents reshape how software is built, and we are adjusting our engineering teams to match. It is an exciting time, and a rare chance to be a strong leader shaping dev tools for the agentic age of coding. Code Plane owns high-stakes, fast-moving products at the center of that shift and needs a tech lead who will set technical and product direction, drive the roadmap from issue to shipped, and keep the team unblocked and moving. You are a generalist by choice. You go where the problem is, backend, AI agent, or frontend, and you are who people come to when it crosses a boundary.

The team’s surface areas include:

  • Batch Changes - our battle-tested execution layer for massive cross-repo changes helps developers automate large-scale code changes across all of an enterprise’s repositories: keeping code up-to-date, fixing critical security issues, and paying down tech debt, with every change tracked through checks and code review until it merges. Teams at Stripe, Uber, and Dropbox, and others use it to drive migrations, refactors, and security fixes.
  • Agentic Batch Changes. The frontier agent for code changes at scale. Building on the foundation of Batch Changes, an outer-loop agent writes  code-modification programs and orchestrates coding agents to roll out changes across the world’s largest codebases, complete with CI feedback loops and post-publish remediation. Anything a single coding agent can do in one repository, Agentic Batch Changes drives across all of them at once, from thousands of repositories to the largest monorepos.
  • Code Monitors and Code Insights - our “signals over code” products. Code Monitors notify and alert based on detected changes in critical code. A leaked secret, a deprecated library, an anti-pattern creeping back in can all turn into alerts. Code Insights turns a codebase into data to back up business decisions and customizable dashboards surfacing critical info about repos so devs can track what matters over time - migration progress, dependency adoption, vulnerability mitigation, code health - with data-driven answers.
  • The src CLI tool - the command-line surface that humans and the agents they employ use to automate Sourcegraph in their workflows. With a natural, scriptable surface, it’s quickly becoming the connective tissue between Sourcegraph and the agent ecosystem.

The features your team ships are used directly by developers and by agents working on their behalf, and you’ll partner closely with Product, Design, Customer Engineering, and adjacent engineering teams to turn customer pain into shipped product.

📅 Within one month, you will…

  • Build relationships across your team and familiarize yourself with its surface area - Batch Changes, Agentic Batch Changes, Code Monitors, Code Insights, src CLI - by running them as a user and an agent would, and form a point of view on the most important product and technical bets to make next.
  • Meet your cross-functional partners in Product, Design, Solutions Engineering, and adjacent engineering teams, and start translating customer feedback into a concrete view of what the team should ship.
  • Join the team’s on-call support rotation.

📅 Within three months, you will…

  • Own a meaningful slice of Agentic Batch Changes end-to-end and ship critical pieces yourself.
  • Be the technical lead for the team’s roadmap, scoping and sequencing projects across the team’s product surfaces, setting clear expectations for ownership, and clarifying priorities.
  • Strengthen the team practices that make their work reliable: how we design, review, test, and deliver.
  • Mentor teammates and continue to build high-trust relationships.

📅 Within six months , you will…

  • Be the recognized technical visionary for Code Plane: to your teammates, and increasingly the wider department, defer to on architecture, quality, and product trade-offs in this space.
  • Be known across the company as the person who deeply understands what customers and agent ecosystems need from Code Plane, and where Sourcegraph should invest next.
  • Be setting the direction of the team’s roadmap with conviction backed by evidence, and mobilizing other engineers to execute on it.

About you

You are a technical leader and software engineer whom people want to follow. You set technical direction, make the hard architecture and tradeoff calls, and keep the team unblocked and moving. You think in terms of customers and outcomes, not tickets. You bring a technical perspective to what the team should and should not take on, and you make the quality bar real: review culture, testing standards, and release practices. You’re scoping with Product and Design and making calls on what to cut so the team can ship. You’re tackling our hardest technical problems hands-on, and also guiding your team to grow.

Agents are one of the most important parts of this job. You have shipped for the agent ecosystem: something that agents or other programs call, with evals behind it, so you know whether a prompt, tool, or model change made things better or worse. You can say precisely where agents fail, because you have encountered it yourself and measured it.

You can take a position, argue it persuasively, and bring people with you, including the ones who started out disagreeing, and you say plainly what would change your mind. You contribute to a collaborative, respectful, async-first culture, and you keep stakeholders informed without being asked.

  • 8+ years of professional software engineering experience
  • Fluency across our stack, backend and frontend: Go, TypeScript, SvelteKit and React, GraphQL APIs, PostgreSQL, and Docker and Kubernetes in a multi-service environment.
  • A strong technical background,  with the ability to provide technical guidance on architecture, code quality, and trade-offs.
  • Proven experience on teams that build end-user product features — ideally developer tools, agent/AI products, CLIs, or other technical product surfaces.
  • Experience partnering closely with Product and Design to take features from idea to shipped, and incorporating customer feedback into the roadmap.
  • You have built something that agents or other programs call, not just something you used an agent to build. MCP servers, agent SDKs, tool surfaces, harnesses, or a CLI designed for non-human callers.
  • Hands-on experience using and reasoning about coding agents (Amp, Claude Code, Cursor, or similar), with strong opinions about what good looks like when humans and agents share a workflow.
  • Async-first communication skills and experience working in a globally distributed remote team.

Strongly preferred

  • You have orchestrated autonomous work at scale and built the machinery that survives it: failures that resume rather than restart, costs that stay inside a budget, and a blast radius you chose deliberately rather than discovered.
  • Experience owning a service you did not build, with paying customers on a schema you inherited, and making it better without a rewrite.
  • A track record of raising the bar for other engineers, such as a testing or code review standard or working with AI assistants
  • You are comfortable making architectural decisions where no established pattern exists, and equally comfortable revising them when the evidence turns, without needing to have been right the first time.
  • You can decide what good looks like for a system that has no single correct answer, build the measurement that proves it, and hold that standard as the ground shifts underneath you.
  • You can sit with a customer, listen past the symptom, and come back with a technical bet the team can build.
  • Experience as a Team or Tech Lead on a software engineering team.

Level

📊 This job is an IC5.  You can read more about our job leveling philosophy in our Handbook.

Compensation

💸 We pay above-market salaries because we want to hire exceptional people who can focus on building great products, not worrying about paying bills. As an open and transparent company, our compensation philosophy and pay bands are visible to every Sourcegraph teammate, and we strive to make our approach equitable, explainable, and competitive.

Your base salary is determined by the IC5 pay band for your location zone (1-4). Our pay bands are informed by market data and designed to ensure competitive compensation wherever you live. During the recruiting process, we’ll discuss the range applicable to you based on job level, relevant skills, experience, qualifications, and location zone.

💰 The starting salary for the IC5 pay band in each zone is:

  • Zone 2: $192,000 USD
  • Zone 3: $144,000 USD

📈 In addition to competitive cash compensation, we offer meaningful equity (because when Sourcegraph succeeds, we want you to succeed, too) and generous perks & benefits.

Interview process

Below is the interview process you can expect for this role (you can read more about the types of interviews in our Handbook). It may look like a lot of steps, but rest assured that we move quickly and the steps are designed to help you get the information needed to determine if we’re the right fit for you… Interviewing is a two-way street, after all!

We expect the interview process to take 4.75 hours in total.

👋 Introduction Stage - we have initial conversations to get to know you better…

  • [30m] Recruiter Screen
  • [60m] Hiring Manager Screen / Resume Deep Dive

🧑‍💻 Team Interview Stage - we then delve into your experience in more depth and introduce you to members of the team, including cross-functional partners…

  • [60m] System Design
  • [60m] Harden a code research agent
  • [60m] Cross-functional team collaboration / Values

🎉 Final Interview Stage- we move you to our final round, where you gain a better understanding of our business and values holistically…

  • [30] Leadership
  • We check references and conduct your background check

Please note - you are welcome to request additional conversations with anyone you would like to meet, but didn’t get to meet during the interview process.

Learn more about us

You can learn more about what it is like to work at Sourcegraph by reading our handbook.

We are an ambitious team who are collectively working hard to build the most influential company in the world. You can read more about our culture, competitive compensation and benefits here.

Sourcegraph is an equal opportunity workplace; we welcome people from all backgrounds.

Sourcegraph participates in E-Verify for U.S. Employees.

Read the full description
Engineer Staff Software Engineer, Model Based at Merlin

Designs and develops flight-critical autonomy algorithms using model-based design tools, Simulink, and DO-178C compliant processes for aerospace systems.

Lead Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

About Merlin:

Merlin (NASDAQ: MRLN) is a publicly traded aerospace and defense company building a non-human pilot to deliver full-stack autonomy for any aircraft from takeoff to touchdown. The Merlin Pilot autonomy system powers a growing range of aircraft and mission profiles and has been proven through hundreds of autonomous flights from Merlin’s global flight test facilities, including Kerikeri, New Zealand; Quonset Point, Rhode Island; and soon, Bedford, Massachusetts. Headquartered in Boston, Merlin is expanding its organization to accelerate the development and deployment of its autonomy platform, helping customers solve some of aviation’s most pressing challenges, from pilot shortages to improving flight safety. Backed by some of the world’s leading investors prior to its public listing, Merlin continues to advance the certification and commercialization of autonomous flight across commercial and defense aviation.

About You:

We are seeking a Staff Software Engineer to design, implement, test, and certify flight-critical autonomy algorithms for next-generation aerospace systems. In this role, you will develop model-based flight software using MathWorks tools and support the full lifecycle of DO-178C compliant development.

Responsibilities:

  • Design and develop flight-critical software using Simulink, Stateflow, and related MathWorks tools for model-based design.
  • Define software architecture, modeling standards, and development workflows aligned with DO-178C and DO-331.
  • Create, maintain and review software requirements, models and auto-generated code.
  • Ensure robustness and traceability through requirements-based design, verification, and certification artifact production.
  • Collaborate with engineers from cross functional groups such as systems, safety, hardware, flight controls, and flight test to ensure product and program level needs are met.
  • Support integration into CI pipelines, including model checks, code generation, static analysis, and automated verification.
  • Contribute to planning and execution of SOI audits and certification reviews.
  • Create and maintain comprehensive documentation for software requirements, architecture and design decisions
  • Support hardware-in-the-loop (HIL), processor-in-the-loop (PIL), and flight testing activities.

Qualifications:

  • Bachelor’s or Master’s degree in Electrical Engineering, Aerospace Engineering, Computer Engineering, Computer Science, or related fields.
  • 10+ years of experience developing embedded or safety-critical software.
  • Extensive experience with Simulink, Stateflow and Embedded Coder for safety critical software development.
  • Experience with Simulink Check, Simulink Code Inspector, Simulink Test and Polyspace Bug Finder
  • Strong experience with requirements management, including authoring high-quality software requirements, maintaining traceability, and using tools such as DOORS, Jama, or Polarion.
  • Working knowledge of DO-178C, including hands-on experience with DO-331.
  • Experience with CI/CD environments and automated model/code quality checks.
  • Experience developing embedded flight software using C/C++ and integrating auto-generated code with manual code
  • Experience performing HIL testing, automated test execution, troubleshooting integration issues and analysis of flight test data.
  • Experience with MATLAB scripting, tool automation, and test automation

$200,000 - $265,000 a year

The compensation range provided is reflective of base salary only, and is a good-faith estimate based on a wide range of factors. The actual offer will be determined by a variety of factors including the candidate’s qualifications, skills, experience, education, and training. Highly competitive equity grants are considered part of Merlin’s total compensation package. Additionally, Merlin offers top-tier benefits for full-time employees.

This position is based on-site at Merlin HQ in Boston, MA.

Once you’re here, you’ll enjoy a variety of on-site perks designed to make your workday enjoyable and convenient. These include catered lunches featuring a rotating menu of delicious options, an assortment of snacks to keep you fueled throughout the day, and a selection of beverages, including coffee, tea, and other drinks, to keep you refreshed.

Our goal is to create an environment where you can thrive both professionally and personally

Merlin Labs offers an innovative, entrepreneurial, and team-focused startup environment. We also offer a top-notch benefits package (health, dental, life, unlimited vacation, and 401k with match) and work/life integration. Being part of the Merlin team allows you to become part of a small team that supports professional development while working together to achieve our mission.

Merlin Labs is an equal opportunity employer and values diversity. We do not discriminate on the basis of race, religion, color, national origin, genetic information, sex (including pregnancy), gender, gender identity and expression, sexual orientation, age, marital status, military service or obligation or disability status, or any other characteristic protected by law. All job offers are contingent upon the candidate passing background and reference checks.

At this time, we are unable to provide visa sponsorship or consider candidates who require visa transfers. Applicants must be authorized to work in the United States without the need for visa sponsorship now or in the future.

In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification form upon hire.

If you require reasonable accommodation in completing an application, interviewing, completing any pre-employment testing, or otherwise participating in the employee selection process, please direct your inquiries to: [email protected]

Merlin Labs does not accept unsolicited resumes from any source other than directly from candidates.

Read the full description
Engineer Tech Lead - Code Plane [IC5] at Sourcegraph

Tech Lead directs engineering strategy and architecture for Sourcegraph's Code Plane products that enable developers and AI agents to navigate and modify large codebases at scale.

Lead Remote Posted about 19 hours ago RemoteFirstJobs Product
What this role involves

Who we are

Our mission is to bring clarity and control to the world’s most complex codebases. AI is accelerating code creation, but the infrastructure to understand, oversee, and evolve that code hasn’t kept pace. Sourcegraph gives engineering organizations full visibility across their systems, precise context for their agents, and the ability to execute coordinated code changes at scale. As agentic development becomes the dominant engineering paradigm, we provide the context layer teams need to take control of their codebase.

With Code Search, Deep Search, MCP, and Agentic Batch Changes, we deliver on that mission today - giving engineering teams and their AI tools the cross-repo context to navigate massive codebases with confidence, and the ability to make changes across hundreds of repositories at once.

Companies like Stripe, Reddit, and Leidos rely on Sourcegraph to ship faster and with higher quality. We’re backed by a16z, Sequoia, and Redpoint, and proud to operate as a globally distributed team that values high agency, direct communication, and customer love.

If you want to build the infrastructure that lets every engineering team - and every agent they deploy - operate on their codebase with confidence, join us.

Hours & location

🌎 While we hire almost anywhere in the world, we have a preference for someone to reside in the following locations for this role. However, if you feel qualified, we welcome you to apply regardless of location. No matter what, working hours must overlap with CEST for at least 20 hours/week.

Preferred locations:

  • Europe
  • EST

Why this job is exciting

The Code Plane team owns the Sourcegraph products that help developers - and the agents working on their behalf - take action on code, at scale, across the world’s largest codebases. Think “data plane” and “control plane” for enterprise code: the surfaces where intent becomes code changes across hundreds (or more!) of repositories.

This is a team that builds products that use AI and products for AI. The roadmap, the architecture, and the day-to-day technical decisions all hinge on a solid experience with and a clear-eyed view of what AI models and agents are good at, where they fail, and how to design products around them. A strong handle on and opinion of AI, formed from real-world, hands-on use, not just observation, is a hard requirement for this role. If you’re excited to ship for the AI agent ecosystem rather than simply watch from the sidelines, you’ll find a lot to love here.

We’re hiring you to be the technical leader around whom this team rallies. Our product surfaces are evolving drastically as agents reshape how software is built, and we are adjusting our engineering teams to match. It is an exciting time, and a rare chance to be a strong leader shaping dev tools for the agentic age of coding. Code Plane owns high-stakes, fast-moving products at the center of that shift and needs a tech lead who will set technical and product direction, drive the roadmap from issue to shipped, and keep the team unblocked and moving. You are a generalist by choice. You go where the problem is, backend, AI agent, or frontend, and you are who people come to when it crosses a boundary.

The team’s surface areas include:

  • Batch Changes - our battle-tested execution layer for massive cross-repo changes helps developers automate large-scale code changes across all of an enterprise’s repositories: keeping code up-to-date, fixing critical security issues, and paying down tech debt, with every change tracked through checks and code review until it merges. Teams at Stripe, Uber, and Dropbox, and others use it to drive migrations, refactors, and security fixes.
  • Agentic Batch Changes. The frontier agent for code changes at scale. Building on the foundation of Batch Changes, an outer-loop agent writes  code-modification programs and orchestrates coding agents to roll out changes across the world’s largest codebases, complete with CI feedback loops and post-publish remediation. Anything a single coding agent can do in one repository, Agentic Batch Changes drives across all of them at once, from thousands of repositories to the largest monorepos.
  • Code Monitors and Code Insights - our “signals over code” products. Code Monitors notify and alert based on detected changes in critical code. A leaked secret, a deprecated library, an anti-pattern creeping back in can all turn into alerts. Code Insights turns a codebase into data to back up business decisions and customizable dashboards surfacing critical info about repos so devs can track what matters over time - migration progress, dependency adoption, vulnerability mitigation, code health - with data-driven answers.
  • The src CLI tool - the command-line surface that humans and the agents they employ use to automate Sourcegraph in their workflows. With a natural, scriptable surface, it’s quickly becoming the connective tissue between Sourcegraph and the agent ecosystem.

The features your team ships are used directly by developers and by agents working on their behalf, and you’ll partner closely with Product, Design, Customer Engineering, and adjacent engineering teams to turn customer pain into shipped product.

📅 Within one month, you will…

  • Build relationships across your team and familiarize yourself with its surface area - Batch Changes, Agentic Batch Changes, Code Monitors, Code Insights, src CLI - by running them as a user and an agent would, and form a point of view on the most important product and technical bets to make next.
  • Meet your cross-functional partners in Product, Design, Solutions Engineering, and adjacent engineering teams, and start translating customer feedback into a concrete view of what the team should ship.
  • Join the team’s on-call support rotation.

📅 Within three months, you will…

  • Own a meaningful slice of Agentic Batch Changes end-to-end and ship critical pieces yourself.
  • Be the technical lead for the team’s roadmap, scoping and sequencing projects across the team’s product surfaces, setting clear expectations for ownership, and clarifying priorities.
  • Strengthen the team practices that make their work reliable: how we design, review, test, and deliver.
  • Mentor teammates and continue to build high-trust relationships.

📅 Within six months , you will…

  • Be the recognized technical visionary for Code Plane: to your teammates, and increasingly the wider department, defer to on architecture, quality, and product trade-offs in this space.
  • Be known across the company as the person who deeply understands what customers and agent ecosystems need from Code Plane, and where Sourcegraph should invest next.
  • Be setting the direction of the team’s roadmap with conviction backed by evidence, and mobilizing other engineers to execute on it.

About you

You are a technical leader and software engineer whom people want to follow. You set technical direction, make the hard architecture and tradeoff calls, and keep the team unblocked and moving. You think in terms of customers and outcomes, not tickets. You bring a technical perspective to what the team should and should not take on, and you make the quality bar real: review culture, testing standards, and release practices. You’re scoping with Product and Design and making calls on what to cut so the team can ship. You’re tackling our hardest technical problems hands-on, and also guiding your team to grow.

Agents are one of the most important parts of this job. You have shipped for the agent ecosystem: something that agents or other programs call, with evals behind it, so you know whether a prompt, tool, or model change made things better or worse. You can say precisely where agents fail, because you have encountered it yourself and measured it.

You can take a position, argue it persuasively, and bring people with you, including the ones who started out disagreeing, and you say plainly what would change your mind. You contribute to a collaborative, respectful, async-first culture, and you keep stakeholders informed without being asked.

  • 8+ years of professional software engineering experience
  • Fluency across our stack, backend and frontend: Go, TypeScript, SvelteKit and React, GraphQL APIs, PostgreSQL, and Docker and Kubernetes in a multi-service environment.
  • A strong technical background,  with the ability to provide technical guidance on architecture, code quality, and trade-offs.
  • Proven experience on teams that build end-user product features — ideally developer tools, agent/AI products, CLIs, or other technical product surfaces.
  • Experience partnering closely with Product and Design to take features from idea to shipped, and incorporating customer feedback into the roadmap.
  • You have built something that agents or other programs call, not just something you used an agent to build. MCP servers, agent SDKs, tool surfaces, harnesses, or a CLI designed for non-human callers.
  • Hands-on experience using and reasoning about coding agents (Amp, Claude Code, Cursor, or similar), with strong opinions about what good looks like when humans and agents share a workflow.
  • Async-first communication skills and experience working in a globally distributed remote team.

Strongly preferred

  • You have orchestrated autonomous work at scale and built the machinery that survives it: failures that resume rather than restart, costs that stay inside a budget, and a blast radius you chose deliberately rather than discovered.
  • Experience owning a service you did not build, with paying customers on a schema you inherited, and making it better without a rewrite.
  • A track record of raising the bar for other engineers, such as a testing or code review standard or working with AI assistants
  • You are comfortable making architectural decisions where no established pattern exists, and equally comfortable revising them when the evidence turns, without needing to have been right the first time.
  • You can decide what good looks like for a system that has no single correct answer, build the measurement that proves it, and hold that standard as the ground shifts underneath you.
  • You can sit with a customer, listen past the symptom, and come back with a technical bet the team can build.
  • Experience as a Team or Tech Lead on a software engineering team.

Level

📊 This job is an IC5.  You can read more about our job leveling philosophy in our Handbook.

Compensation

💸 We pay above-market salaries because we want to hire exceptional people who can focus on building great products, not worrying about paying bills. As an open and transparent company, our compensation philosophy and pay bands are visible to every Sourcegraph teammate, and we strive to make our approach equitable, explainable, and competitive.

Your base salary is determined by the IC5 pay band for your location zone (1-4). Our pay bands are informed by market data and designed to ensure competitive compensation wherever you live. During the recruiting process, we’ll discuss the range applicable to you based on job level, relevant skills, experience, qualifications, and location zone.

💰 The starting salary for the IC5 pay band in each zone is:

  • Zone 2: $192,000 USD
  • Zone 3: $144,000 USD

📈 In addition to competitive cash compensation, we offer meaningful equity (because when Sourcegraph succeeds, we want you to succeed, too) and generous perks & benefits.

Interview process

Below is the interview process you can expect for this role (you can read more about the types of interviews in our Handbook). It may look like a lot of steps, but rest assured that we move quickly and the steps are designed to help you get the information needed to determine if we’re the right fit for you… Interviewing is a two-way street, after all!

We expect the interview process to take 4.75 hours in total.

👋 Introduction Stage - we have initial conversations to get to know you better…

  • [30m] Recruiter Screen
  • [60m] Hiring Manager Screen / Resume Deep Dive

🧑‍💻 Team Interview Stage - we then delve into your experience in more depth and introduce you to members of the team, including cross-functional partners…

  • [60m] System Design
  • [60m] Harden a code research agent
  • [60m] Cross-functional team collaboration / Values

🎉 Final Interview Stage- we move you to our final round, where you gain a better understanding of our business and values holistically…

  • [30] Leadership
  • We check references and conduct your background check

Please note - you are welcome to request additional conversations with anyone you would like to meet, but didn’t get to meet during the interview process.

Learn more about us

You can learn more about what it is like to work at Sourcegraph by reading our handbook.

We are an ambitious team who are collectively working hard to build the most influential company in the world. You can read more about our culture, competitive compensation and benefits here.

Sourcegraph is an equal opportunity workplace; we welcome people from all backgrounds.

Sourcegraph participates in E-Verify for U.S. Employees.

Read the full description
Engineer Legion: Director of Production Engineering

Leads DevOps and SRE teams to build reliable, scalable, and secure AWS-based production infrastructure while spending 20-30% time on hands-on architecture and tooling work.

Lead Remote Posted 1 day ago We Work Remotely — Programming
What this role involves

Headquarters: Remote, United States

Director of Production Engineering 

Remote, United States

About this Position

Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps & SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance.

This is a hands-on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross-team execution.

Responsibilities

  • Hire and build a globally-distributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.
  • Own the reliability and infrastructure roadmap for our AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.
  • Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.
  • Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.
  • Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution.
  • Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and DevOps processes (e.g., automated investigation, remediation, and PR-generation pipelines).
  • Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.
  • Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.
  • Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit-readiness.
  • Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals.
  • Stay current on cloud, DevOps, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.

Required Qualifications

  • 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience.
  • Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).
  • Demonstrated experience running security operations (SecOps) — vulnerability management, incident response, and remediation of production security issues.
  • 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.
  • Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.
  • Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines.
  • Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS).
  • Proven track record partnering cross-functionally with engineering and IT teams to align on infrastructure, tooling, and security standards.
  • Demonstrated experience leading incident management and on-call practices for high-availability production systems.
  • Bachelor's degree in Computer Science, Engineering, or related field required; Master's degree preferred.

Preferred Qualifications

  • Experience with major cloud providers beyond AWS, such as Google Cloud Platform or Oracle Cloud Infrastructure (OCI).
  • Relevant security certifications (e.g., CISSP, AWS Security Specialty, CKS).
  • Experience with compliance frameworks such as SOC 2, ISO 27001, or HIPAA.
  • 3+ years of experience with Kubernetes or other containerization/orchestration platforms at scale.
  • Experience with Kubernetes-native delivery tooling, including Argo Workflows and Helm.
  • 5+ years of experience leading teams in an agile/scrum environment.
  • Experience building or scaling automated investigation and remediation pipelines for production error classes.

COMPENSATION & BENEFITS

Salary Range: Base Salary Range  $220,000 - $265,000 + Bonus + Stock Equity

At Legion, we offer competitive compensation and benefits packages to all employees. As a fully remote employer, pay for positions is determined using local, national, and industry-specific survey data. 

Our posted salary range is done so in good faith based on national data and may be refined for a candidate's region/town/cost of living. We strive to make competitive offers that allow employees room for future growth. Salaries will be based on the applicant’s location, level of experience, education, and specialized knowledge and skills. Additionally, we consider the external market rate, the amount we have budgeted internally, and the internal equity for the same position within the company. 

Benefits include, but are not limited to:

  • $0 monthly premium and other flexible medical, dental, and vision plans effective on the first day of employment
  • 401k plan
  • Discretionary Paid Time Off and Paid Holidays
  • Parental Leave 
  • Equity 
  • Monthly Wellness Reimbursement
  • Monthly Lunch on Legion

ABOUT LEGION

Join Legion's mission to turn hourly jobs into good jobs. We're a remote, mission-driven team seeking exceptional talent to propel this vision. Embrace a culture that's collaborative, fast-paced, and entrepreneurial. With us, you'll grow your skills, work closely with experienced executives, and contribute significantly to our mission.

Legion Technologies delivers the industry’s most innovative workforce management platform. It enables businesses to maximize labor efficiency and employee engagement simultaneously. The award-winning, AI-driven Legion WFM platform is intelligent, automated, and employee-centric. It’s proven to deliver 13x ROI through schedule optimization, reduced attrition, increased productivity, and increased operational efficiency. Legion delivers cutting-edge technology in an easy-to-use platform and mobile app that employees love. 

If you're ready to make an impact and grow your career, Legion is where you belong. Join us in making hourly work rewarding and fulfilling.

BACKGROUND AND OPPORTUNITY 

There are almost 75 million hourly workers in the United States, representing more than half of the entire workforce. Historically, managing hourly employees has been difficult due to high attrition (average of 60%) and high replacement costs (average of $3,200 per employee in retail). The ongoing labor shortage and competition from the gig economy make it more difficult to attract and retain hourly employees. The top reasons hourly employees leave their jobs are a lack of schedule empowerment, poor communication with employers, and an inability to get paid early. Gen Z and the millennial workforce demand gig-like flexibility, modern technology, and compelling work options. Legion’s mission is to turn hourly jobs into good jobs, serving the hourly workers who make up the majority of the US workforce. We believe in empowering employees and helping employers be efficient and innovative by enabling intelligent automation powered by Legion’s Workforce Management platform to optimize labor efficiency and enhance the employee experience simultaneously. Legion WFM was built for the cloud with AI at the core and designed to handle the complexity of modern businesses and meet the needs of today’s hourly employees.  Our team is comprised of dedicated individuals from all backgrounds and experiences, globally distributed across all time zones.

For more information, visit https://legion.co 

EQUAL EMPLOYMENT OPPORTUNITY

Legion Technologies is proud to be an equal-opportunity employer and is committed to maintaining a diverse and inclusive work environment. All qualified applicants will be considered for employment without regard to race, color, religion, sex, age, disability, marital status, familial status, sexual orientation, pregnancy, genetic information, gender identity, gender expression, national origin, ancestry, citizenship status, veteran status, and any other legally protected status under federal, state, or local anti-discrimination laws.

DISABILITY ACCOMMODATION

For individuals with disabilities who need additional assistance at any point in the application and interview process, please email recruiting@legion.co 

We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Legion recruiting email communications will always come from the @legion.co domain. Any outreach claiming to be from Legion via other sources should be ignored. If you are uncertain whether you have been contacted by an official Legion employee, reach out to recruiting@legion.co

 

Legion is an equal opportunity employer. All applicants will be considered for employment without attention to race, religion, color, sex, sexual orientation, gender identity, age, national origin, veteran, disability status, or any other basis covered by appropriate law.

How We Determine What We Pay

As a global employer, Legion determines pay for positions using local, national, and industry-specific survey data. We evaluate external equity and the cost of labor/prevailing wage index in the relative marketplace for jobs directly comparable to jobs within our company. Our posted salary range is based on national data and may be refined for a candidate's region/town/cost of living. For new hires, we strive to make competitive offers allowing the new employee room for future growth. Salaries will be based on the applicant’s location, level of experience, education, and specialized knowledge and skills. Additionally, we consider the external market rate, the amount we have budgeted internally, and internal equity within the company for the same position. An employee/candidate with a stronger skill set will receive higher pay.

Job Applicant Privacy Policy

This Job Applicant Privacy Policy (“Policy”) describes how Legion Technologies, Inc. (“Legion”, “we”, “us” and “our”) collects, uses, and discloses “personal information” as defined under California law from and about job applicants who are residents of California.

This Policy does not apply to our handling of data gathered about you in your role as a user of our consumer-facing services. When you interact with us as in that role, the Legion Privacy Policy applies.

  1. Types of Personal Information We Handle

    We collect, store, and use various types of personal information through the application and recruitment process. We collect such information either directly from you or (where applicable) from another person or entity, such as an employment agency or consultancy, background check provider, or other referral sources. This information includes:

    • Identification and contact information, and related identifiers such as full name, date and place of birth, citizenship and permanent residence, home and business addresses, telephone numbers, email addresses, and such information about your beneficiaries or emergency contacts.
    • Professional or employment-related information, including:
      • Recruitment, employment, or engagement information such as application forms and information included in a resume, cover letter, or otherwise provided through any application or engagement process; and copies of identification documents, such as driver’s licenses, passports, and visas; and background screening results and references.
      • Career information such as job titles; work history; work dates and work locations; information about skills, qualifications, experience, publications, speaking engagements, and preferences; and professional memberships
    • Education Information such as institutions attended, degrees, certifications, training courses, publications, and transcript information.
    • Legally protected classification information such as race, sex/gender, religious/ philosophical beliefs, gender identity/expression, sexual orientation, marital status, military service, nationality, ethnicity, request for family care leave, political opinions, and criminal history.
    • Other information such as any information you voluntarily choose to provide in connection with your job application.
  2. How We Use Personal Information

    We collect, use, share, and store personal information from job applicants for our and our service providers’ business and operational purposes in the recruitment process such as: processing your application, tracking your application through the recruitment process, contacting references with your authorization, conducting background checks you authorize, and making hiring decisions. We will also use job applicant information for internal analysis purposes to understand the applicants who apply and to improve our recruitment process. We may sometimes need to use applicant information for legal purposes, such as in connection with any challenges made to our hiring decisions.

  3. With Whom We Share Personal Information

    We will disclose job applicant personal information to the following types of entities or in the following circumstances (where applicable):

    • Internally: to other Legion personnel involved in the recruiting and hiring process.
    • Vendors: such as technology service providers, travel management providers, human resources suppliers, background check companies, and employment agencies or recruiters, where applicable.
    • Legal Compliance: when required to do so by law, regulation, or court order or in response to a request for assistance by the police or other law enforcement agency.
    • Litigation Purposes: to seek legal advice from our external lawyers or in connection with litigation with a third party.
    • Business Transaction Purposes: in connection with the sale, purchase, or merger.
  4. How to Contact Us About this Policy – If you have any questions about this Policy, please contact privacy@legion.co.

To apply: https://weworkremotely.com/remote-jobs/legion-director-of-production-engineering

Read the full description
Engineer Principal Software Engineer at HubSpot

Designs and builds HubSpot's observability platform for distributed systems and AI agents, setting architecture patterns and telemetry standards across hundreds of microservices.

Lead Posted 2 days ago RemoteFirstJobs Product
What this role involves

POS-5690

About the Team

The Observability team owns the internal platform that gives every HubSpot engineer real visibility into how their systems behave in production. We build and operate the distributed tracing, metrics, alerting, and logging infrastructure that spans hundreds of microservices, billions of daily events, and thousands of engineers who depend on that signal to ship reliably.

We are now investing in the next generation of this platform. As HubSpot deploys AI agents and ML-powered features across the product, the team is building the tracing and telemetry primitives that make it possible to understand, debug, and trust what those systems are doing in production. This is greenfield, technically interesting work at a scale few companies operate at and we are looking for a Principal Engineer to help lead it.

About the Role

We are seeking a Principal Software Engineer to be the technical anchor for HubSpot’s Observability platform. This role sits at the intersection of large-scale distributed systems, developer platform design, and AI observability. A big part of this role is working horizontally across a large engineering org and setting patterns and standards that make it easier for teams to instrument, alert on, and reason about their services. You will also shape how we trace and understand our growing fleet of AI agents and ML systems in production: a technically distinct and increasingly critical problem.

Key Expectations

  • Observability Platform Architecture: Define the patterns and evolution of HubSpot’s core telemetry platform — distributed tracing, metrics, and structured logging — at a scale that spans hundreds of services and billions of daily events. Set the standards for how instrumentation is done across a large, polyglot engineering organization.
  • AI & Agentic Observability: Lead the technical strategy for tracing and understanding AI agents and ML-powered systems in production. Define the primitives, telemetry standards, and debugging workflows that help product engineers understand what their models and agents are doing — and build trust in those systems over time. This is greenfield and consequential work.
  • High-Cardinality, High-Throughput Systems: Architect telemetry pipelines and storage systems that handle high-cardinality data at high throughput without blowing up cost or query latency. Make principled tradeoffs between sampling, fidelity, retention, and developer ergonomics.
  • Hands-on, High-Leverage Builder: Ship production code. Lead design reviews and take high-impact initiatives end-to-end, from prototype to production system at scale. Stay close to the systems you build and be the person who can debug the hardest problems when they surface.
  • Developer Experience & Adoption: Design the instrumentation APIs and libraries that product engineers reach for, making correct observability the path of least resistance. Drive OpenTelemetry adoption across a large, polyglot codebase. Build the tooling that turns raw telemetry into actionable signal for teams operating at speed.
  • Production Intelligence & Reliability Patterns: Define patterns for SLO/SLI design, alerting philosophy, and how teams graduate from reactive to proactive incident response. Push for simplicity in a domain that wants to get complicated, and consistency where tooling can drift across a large organization.
  • Technical Leadership & Influence: Partner with infrastructure, platform, and product engineering teams to understand their signal gaps and close them. Influence technical strategy alongside engineering leadership, translating observability constraints and opportunities into product and operational decisions. Mentor senior engineers and tech leads, driving thoughtful design decisions and capturing learnings from major incidents and large-scale migrations.

What You Bring

  • Platform-Builder Experience: Proven experience building observability or telemetry tooling for internal engineering teams, rather than simply consuming it. You understand how to architect developer platforms that serve thousands of engineers across a large organization, backed by deep operational instincts and hard-earned expertise.
  • Telemetry Systems Depth: Deep expertise navigating trade-offs in telemetry pipeline design across high-cardinality data, dynamic sampling, query latency, retention economics, and data fidelity. Strong technical fluency with OpenTelemetry, distributed tracing engines, metric backends, and large-scale log ingestion infrastructure.
  • Incident Automation & Operational Excellence: Proven track record linking telemetry signals directly to automated operational workflows. You have designed architecture for real-time telemetry triggers that power automated remediation, dynamic runbooks, or AI-assisted root-cause diagnosis across microservices environments.

Why This Role, Why Now

HubSpot is scaling fast — more engineers, more microservices, more AI systems running in production — and the Observability platform is at an inflection point. The foundations are solid, but the next phase requires a different kind of investment: rethinking cardinality economics, making OpenTelemetry the default across a large polyglot org, and building an entirely new layer of AI tracing that doesn’t yet exist.

This Principal Engineer will have a direct line of sight from the architecture they design to the engineering outcomes we measure: incident response times, adoption rates, developer satisfaction, and the trust that product teams place in their production signal. The AI observability layer in particular is greenfield — there is no playbook to follow, which is exactly what makes this the right moment for the right person.

If you want to build the platform that helps thousands of engineers understand what their systems are doing — including systems powered by AI that are genuinely hard to see inside — this is the role.

Pay & Benefits

The cash compensation below includes base salary, on-target commission for employees in eligible roles, and annual bonus targets under HubSpot’s bonus plan for eligible roles. In addition to cash compensation, some roles are eligible to participate in HubSpot’s equity plan to receive restricted stock units (RSUs). Some roles may also be eligible for overtime pay. Individual compensation packages are tailored to your skills, experience, qualifications, and other job-related reasons.

This resource will help guide how we recommend thinking about the range you see. Learn more about HubSpot’s compensation philosophy.

Benefits are also an important piece of your total compensation package. Explore the benefits and perks HubSpot offers to help employees grow better.

At HubSpot, fair compensation practices aren’t just about checking off the box for legal compliance. It’s about living out our value of transparency with our employees, candidates, and community.

Annual Cash Compensation Range:

$313,800—$502,080 USD

At HubSpot, we value both flexibility and connection. Whether you’re a Remote employee or work from the Office, we want you to start your journey here by building strong connections with your team and peers. If you are joining our Engineering team, you will be required to attend a regional HubSpot office for in-person onboarding. If you join our broader Product team, you’ll also attend other in-person events, such as your Product Group Summit and other gatherings, to continue building on those connections.

If you require an accommodation due to travel limitations or other reasons, please inform your recruiter during the hiring process. We are committed to supporting candidates who may need alternative arrangements

About HubSpot

HubSpot (NYSE: HUBS) is an AI-powered customer platform with all the software, integrations, and resources customers need to connect marketing, sales, and service. HubSpot’s connected platform enables businesses to grow faster by focusing on what matters most: customers.

At HubSpot, bold is our baseline. Our employees around the globe move fast, stay customer-obsessed, and win together. Our culture is grounded in four commitments: Solve for the Customer, Be Bold, Learn Fast, Align, Adapt & Go!, and Deliver with HEART. These commitments shape how we work, lead, and grow.

We’re building a company where people can do their best work. We focus on brilliant work, not badge swipes. By combining clarity, ownership, and trust, we create space for big thinking and meaningful progress. And we know that when our employees grow, our customers do too.

Recognized globally for our award-winning culture by Comparably, Glassdoor, Fortune, and more, HubSpot is headquartered in Cambridge, MA, with employees and offices around the world.

Explore more:

  • HubSpot Careers
  • Life at HubSpot on Instagram

If you need accommodations or assistance due to a disability, please reach out to us using this form.

Massachusetts Applicants: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Germany Applicants: (m/f/d) - link to HubSpot’s Career Diversity page here.

India Applicants: link to HubSpot India’s equal opportunity policy here.

HubSpot may use AI to help screen or assess candidates, but all hiring decisions are always human. More information can be found here. By submitting your application, you agree that HubSpot may collect your personal data for recruiting, global organization planning, and related purposes. We may use CLEAR ID Verification during the hiring process to confirm your identity and help maintain a safe, secure, and trusted experience for all candidates. Refer to HubSpot’s Recruiting Privacy Notice for details on data processing and your rights.

Read the full description
Engineer Staff Site Reliability Engineer at Replit

Staff SRE architecting observability solutions, defining reliability standards, leading incident response, and automating infrastructure at scale.

Lead Posted 2 days ago RemoteFirstJobs Product
What this role involves

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the role:

Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit’s infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability.

We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure’s reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit.

You Will:

  • Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection.

  • Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed.

  • Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution. Conduct thorough, blameless post-mortems and drive the implementation of preventative measures. Develop and refine runbooks and build automation to reduce Mean Time To Recovery (MTTR).

  • Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios.

  • Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments, with a deep focus on Kubernetes, Docker, and GCP. Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions.

  • Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack. Use your findings to design and implement long-term fixes that make our systems and products more robust, operable, and easier to diagnose.

  • Provide Staff-Level Guidance: Review feature and system designs from across the company, acting as a key owner for the reliability, scalability, security, and operational integrity of those designs.

  • Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems, making reliability a core value of the Replit engineering culture.

  • Build and Integrate: Write high-quality, well-tested code in Python or Go to meet the needs of your customers, whether it’s building new internal tools or integrating with third-party vendors.

Required Skills and Experience:

  • 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).

  • Strong programming skills in languages like Python or Go. You write high-quality, well-tested code.

  • Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture.

  • Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.

  • Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing).

  • Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure.

  • Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.

  • Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices.

  • Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels.

  • A willingness to dive into understanding, debugging, and improving any layer of the stack.

  • You’re passionate about making software creation accessible and empowering the next generation of builders.

Bonus Points:

  • Deep experience with Google Cloud Platform (GCP) services and tools.

  • Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry).

  • Experience designing and building reliable systems capable of handling high throughput and low latency.

  • Significant experience with Go and Terraform.

  • Familiarity with working in rapid-growth, startup environments.

  • Experience writing company-facing blog posts and training materials.

Full-Time Employee Benefits Include:

💰 Competitive Salary & Equity

💹 401(k) Program with a 4% match ( US Only)

⚕️ Health, Dental, Vision and Life Insurance

🩼 Short Term and Long Term Disability

🚼 Paid Parental, Medical, Caregiver Leave

🏝 Flexible Time Off (FTO) + Holidays

🚗 Commuter Benefits ( In-Office & US Only)

📱 Monthly Wellness Stipend

🧑‍💻 Autonomous Work Environment

🖥 In Office Set-Up Reimbursement ( In-Office Only)

🚀 Quarterly Team Gatherings

☕ In Office Amenities ( In-Office Only)

Want to learn more about what we are up to?

  • Self-driving Company

  • Replit Agent at Scale

  • AI Adoption

  • Build Open-Source Apps

Interviewing + Culture at Replit

  • Operating Principles

  • Reasons not to work at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Read the full description
Engineer Principal Software Engineer at HubSpot

Designs and leads HubSpot's observability platform architecture, including distributed tracing, metrics, logging, and AI/ML system telemetry at scale across hundreds of microservices.

Lead Posted 2 days ago RemoteFirstJobs Product
What this role involves

POS-5690

About the Team

The Observability team owns the internal platform that gives every HubSpot engineer real visibility into how their systems behave in production. We build and operate the distributed tracing, metrics, alerting, and logging infrastructure that spans hundreds of microservices, billions of daily events, and thousands of engineers who depend on that signal to ship reliably.

We are now investing in the next generation of this platform. As HubSpot deploys AI agents and ML-powered features across the product, the team is building the tracing and telemetry primitives that make it possible to understand, debug, and trust what those systems are doing in production. This is greenfield, technically interesting work at a scale few companies operate at and we are looking for a Principal Engineer to help lead it.

About the Role

We are seeking a Principal Software Engineer to be the technical anchor for HubSpot’s Observability platform. This role sits at the intersection of large-scale distributed systems, developer platform design, and AI observability. A big part of this role is working horizontally across a large engineering org and setting patterns and standards that make it easier for teams to instrument, alert on, and reason about their services. You will also shape how we trace and understand our growing fleet of AI agents and ML systems in production: a technically distinct and increasingly critical problem.

Key Expectations

  • Observability Platform Architecture: Define the patterns and evolution of HubSpot’s core telemetry platform — distributed tracing, metrics, and structured logging — at a scale that spans hundreds of services and billions of daily events. Set the standards for how instrumentation is done across a large, polyglot engineering organization.
  • AI & Agentic Observability: Lead the technical strategy for tracing and understanding AI agents and ML-powered systems in production. Define the primitives, telemetry standards, and debugging workflows that help product engineers understand what their models and agents are doing — and build trust in those systems over time. This is greenfield and consequential work.
  • High-Cardinality, High-Throughput Systems: Architect telemetry pipelines and storage systems that handle high-cardinality data at high throughput without blowing up cost or query latency. Make principled tradeoffs between sampling, fidelity, retention, and developer ergonomics.
  • Hands-on, High-Leverage Builder: Ship production code. Lead design reviews and take high-impact initiatives end-to-end, from prototype to production system at scale. Stay close to the systems you build and be the person who can debug the hardest problems when they surface.
  • Developer Experience & Adoption: Design the instrumentation APIs and libraries that product engineers reach for, making correct observability the path of least resistance. Drive OpenTelemetry adoption across a large, polyglot codebase. Build the tooling that turns raw telemetry into actionable signal for teams operating at speed.
  • Production Intelligence & Reliability Patterns: Define patterns for SLO/SLI design, alerting philosophy, and how teams graduate from reactive to proactive incident response. Push for simplicity in a domain that wants to get complicated, and consistency where tooling can drift across a large organization.
  • Technical Leadership & Influence: Partner with infrastructure, platform, and product engineering teams to understand their signal gaps and close them. Influence technical strategy alongside engineering leadership, translating observability constraints and opportunities into product and operational decisions. Mentor senior engineers and tech leads, driving thoughtful design decisions and capturing learnings from major incidents and large-scale migrations.

What You Bring

  • Platform-Builder Experience: Proven experience building observability or telemetry tooling for internal engineering teams, rather than simply consuming it. You understand how to architect developer platforms that serve thousands of engineers across a large organization, backed by deep operational instincts and hard-earned expertise.
  • Telemetry Systems Depth: Deep expertise navigating trade-offs in telemetry pipeline design across high-cardinality data, dynamic sampling, query latency, retention economics, and data fidelity. Strong technical fluency with OpenTelemetry, distributed tracing engines, metric backends, and large-scale log ingestion infrastructure.
  • Incident Automation & Operational Excellence: Proven track record linking telemetry signals directly to automated operational workflows. You have designed architecture for real-time telemetry triggers that power automated remediation, dynamic runbooks, or AI-assisted root-cause diagnosis across microservices environments.

Why This Role, Why Now

HubSpot is scaling fast — more engineers, more microservices, more AI systems running in production — and the Observability platform is at an inflection point. The foundations are solid, but the next phase requires a different kind of investment: rethinking cardinality economics, making OpenTelemetry the default across a large polyglot org, and building an entirely new layer of AI tracing that doesn’t yet exist.

This Principal Engineer will have a direct line of sight from the architecture they design to the engineering outcomes we measure: incident response times, adoption rates, developer satisfaction, and the trust that product teams place in their production signal. The AI observability layer in particular is greenfield — there is no playbook to follow, which is exactly what makes this the right moment for the right person.

If you want to build the platform that helps thousands of engineers understand what their systems are doing — including systems powered by AI that are genuinely hard to see inside — this is the role.

Pay & Benefits

The cash compensation below includes base salary, on-target commission for employees in eligible roles, and annual bonus targets under HubSpot’s bonus plan for eligible roles. In addition to cash compensation, some roles are eligible to participate in HubSpot’s equity plan to receive restricted stock units (RSUs). Some roles may also be eligible for overtime pay. Individual compensation packages are tailored to your skills, experience, qualifications, and other job-related reasons.

This resource will help guide how we recommend thinking about the range you see. Learn more about HubSpot’s compensation philosophy.

Benefits are also an important piece of your total compensation package. Explore the benefits and perks HubSpot offers to help employees grow better.

At HubSpot, fair compensation practices aren’t just about checking off the box for legal compliance. It’s about living out our value of transparency with our employees, candidates, and community.

Annual Cash Compensation Range:

$313,800—$502,080 USD

At HubSpot, we value both flexibility and connection. Whether you’re a Remote employee or work from the Office, we want you to start your journey here by building strong connections with your team and peers. If you are joining our Engineering team, you will be required to attend a regional HubSpot office for in-person onboarding. If you join our broader Product team, you’ll also attend other in-person events, such as your Product Group Summit and other gatherings, to continue building on those connections.

If you require an accommodation due to travel limitations or other reasons, please inform your recruiter during the hiring process. We are committed to supporting candidates who may need alternative arrangements

About HubSpot

HubSpot (NYSE: HUBS) is an AI-powered customer platform with all the software, integrations, and resources customers need to connect marketing, sales, and service. HubSpot’s connected platform enables businesses to grow faster by focusing on what matters most: customers.

At HubSpot, bold is our baseline. Our employees around the globe move fast, stay customer-obsessed, and win together. Our culture is grounded in four commitments: Solve for the Customer, Be Bold, Learn Fast, Align, Adapt & Go!, and Deliver with HEART. These commitments shape how we work, lead, and grow.

We’re building a company where people can do their best work. We focus on brilliant work, not badge swipes. By combining clarity, ownership, and trust, we create space for big thinking and meaningful progress. And we know that when our employees grow, our customers do too.

Recognized globally for our award-winning culture by Comparably, Glassdoor, Fortune, and more, HubSpot is headquartered in Cambridge, MA, with employees and offices around the world.

Explore more:

  • HubSpot Careers
  • Life at HubSpot on Instagram

If you need accommodations or assistance due to a disability, please reach out to us using this form.

Massachusetts Applicants: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Germany Applicants: (m/f/d) - link to HubSpot’s Career Diversity page here.

India Applicants: link to HubSpot India’s equal opportunity policy here.

HubSpot may use AI to help screen or assess candidates, but all hiring decisions are always human. More information can be found here. By submitting your application, you agree that HubSpot may collect your personal data for recruiting, global organization planning, and related purposes. We may use CLEAR ID Verification during the hiring process to confirm your identity and help maintain a safe, secure, and trusted experience for all candidates. Refer to HubSpot’s Recruiting Privacy Notice for details on data processing and your rights.

Read the full description
Engineer Senior Staff Software Engineer, Partner Integrations at Flex

Senior Staff Software Engineer oversees technical roadmap for partner integrations, building APIs and SDKs that enable external partners to connect to Flex's rent payment platform.

Lead Posted 2 days ago RemoteFirstJobs Product
What this role involves

Flex is a growth-stage, NYC headquartered FinTech company that is creating the best rent payment experience. It’s hard to believe that it’s 2026 and paying rent on time is expensive, inflexible, and difficult. We’re here to change that! Flex enables our users to pay rent throughout the month on a schedule that better fits their finances and budget. Our mission is to empower as many renters as possible with flexibility over their most significant recurring expense. After deliberately keeping a stealth profile as we built up unprecedented investor support and an enthusiastic user base, we are looking for motivated individuals to help us keep our mission growing. Will you be a part of the team?

About Our Opportunity

Flex exists to make paying rent work the way real life actually works — smoothing out cash flow, helping renters avoid late fees, and letting them build credit instead of falling behind. As a Senior Staff Software Engineer, you’ll oversee the technical roadmap for the team, working across APIs, SDKs, and Web experiences. You will work with teams across the organization including Engineering, Product, Design, Infrastructure, Sales, Partner and Customer Success to ensure that the technical strategy meets our goals. We expect you to be hands-on and execute work as an individual, and build products that allow for flexibility as we evolve our product offerings.

About Our Team!

We have some really exciting teams who are looking for an amazing engineer like you! Check them out!

  • Partner Integrations - Partner integrations work to standardize APIs, tools, and processes that let external partners connect to Flex for data sharing and payments across entry points like Flex Anywhere

Who thrives here?

  • A Doer— you get things done efficiently, and you’re comfortable rolling up your sleeves when something’s blocking the team.
  • An Owner — you take full accountability for what you build, from design through production, and you always have a plan B.
  • Someone Collaborative— you bring people together to pressure-test ideas rather than working in a silo.
  • Someone with Precision — your code and your communication are both clear, because ideas only matter once others can act on them.
  • Someone Resilient— systems break and priorities shift, and you adapt without losing momentum.
  • Someone Humble— you care more about getting to the right answer than being the one who’s right.

Qualifications:

  • Bachelor or above degree in Computer Science or a related field
  • Minimum of 10 years experience in software engineering, with at least 3 years of technical leadership experience in a hands-on capacity
  • Ability to work on a globally-distributed team with a high degree of ownership
  • Experience leading the delivery of multiple highly impactful products end to end, on time with a high quality bar
  • Experience working with technical and non-technical stakeholders, successfully aligning and setting expectations on scope and delivery
  • Ability to drive yourself and your team to bring quality and consistency to their code and architecture, without compromising velocity
  • Ability to grow in a fast-paced and dynamic environment that will challenge you to always bring your best
  • Experience in designing and developing solutions that maximize ROI and lead to significant business impact
  • Experience working in FinTech and familiar with major payment rails
  • Experience working with Sales stakeholders and/or building products for Sales professionals
  • Experience building integrations with external partners & managing relationships with external stakeholders
  • Experience working on AWS cloud based applications

Compensation

Flex takes a market-based approach to pay, and compensation may vary depending on your primary work location. Work locations are categorized into one of three tiers based on a cost of labor index for that geographic area. The successful candidate’s starting pay will be commensurate with their experience, qualifications, and Flex’s internal leveling guidelines and benchmarks.

Tier 1 (NYC/Bay Area, Los Angeles, Seattle)

$240,000—$300,000 USD

Tier 2 (Austin, Washington D.C. Philadelphia, San Diego, Chicago, Atlanta)

$216,000—$270,000 USD

Tier 3 (Salt Lake City, all other USA cities)

$204,000—$255,000 USD

Life at Flex

We understand that it takes a diverse team of highly intelligent, curious, determined, empathetic, and self aware people to grow a successful company. Our HQ is located in New York City, but we have employees located throughout the US, Australia, Canada and South America. We are growing quickly, but deliberately, with a focus on building an inclusive culture. Our dynamic team has incredible perspectives to share, just as we know you do, and we take great pride in being an equal opportunity workplace.

Offices

Roles posted in New York, San Francisco, and Salt Lake City are hybrid positions with on-site expectations of 2-3 days per week in our local offices. For candidates outside of these areas, you may be eligible for our relocation assistance program.

Benefits

For full-time U.S. employees we offer:

  • Competitive medical, dental, and vision
  • Company equity
  • 401(k) plan with company match
  • Unlimited paid time off + 13 company paid holidays
  • Parental leave
  • Free Flex subscription

For full-time non-U.S. employees, we offer:

  • Competitive compensation + company equity
  • Unlimited PTO
Read the full description
Engineer Staff Software Backend Engineer, Tech Lead – Foundations

Technically leads the Foundations platform team, building shared backend infrastructure and primitives that other engineering teams depend on.

Lead Posted 2 days ago Jobicy AI
What this role involves
NetBox Labs seeks a Staff Software Backend Engineer to technically lead our Foundations team. Foundations is the platform team that builds shared platform primitives that other engineering teams depend on...
Read the full description
Engineer Director – Forward Deployed Engineering at Re:Build Manufacturing

Leads a team of forward-deployed engineers building full-stack digital solutions across manufacturing sites, managing technical strategy, delivery, and people development.

Lead Posted 3 days ago RemoteFirstJobs Product
What this role involves

About Re:Build

At Re:Build, our mission is to ensure the next generation of important products are made, at scale, in America. We are laying the foundation for a better future for our customers, employees, and communities by revitalizing America’s manufacturing base and creating meaningful jobs across the country, including in historically deindustrialized regions.

We operate an advanced, end-to-end manufacturing platform that partners with industrial companies and innovators to take products from first concept to full-scale production in critical verticals including aerospace and defense, electrification, medical, energy and environment, and robotics and automation.

The way we operate is as important as the work we do. It’s guided by The Re:Build Way, 16 principles that shape how we collaborate with each other, partner with our customers and vendors, and contribute to the communities where we operate. (link to The Re:Build Way principles )

Who we are looking for

The Director, Forward Deployed Engineering is a hands-on leadership role in Re:Build Manufacturing’s Digital Innovation Group (DIG). This role reflects the organization’s belief that digital innovation is key to scaling as a premier industrial company. In this position, the Director is responsible for hiring, leading, and developing a team of Forward Deployed Engineers (FDEs) who are embedded across Re:Build sites and Resource Center functions, implementing full-stack digital solutions that solve real-world business problems and drive measurable business impact. Success in the role depends on the ability to build trust, ensure alignment, and partner effectively with internal stakeholders throughout the entire product lifecycle. The ideal candidate combines a strong software engineering background with proven experience managing software delivery and solution engineering that directly serves customers. They bring deep technical credibility, strong people leadership skills, and the judgment needed to keep engagements well prioritized and delivered quickly.

What you get to do

  • Partner with portfolio Manufacturing, Engineering Services, and Functional teams to identify difficulties, operational inefficiencies, and opportunities for digital solutions
  • Evaluate and prioritize digital solution opportunities based on business impact, feasibility, and strategic alignment
  • Guide FDEs through sophisticated technical decisions, feature development, fixing issues, code reviews, and delivery execution across collaborator engagement
  • Define and implement engineering standards and coding practices across FDE projects to ensure scalability, maintainability, and security of delivered solutions
  • Build the repeatable delivery machine of reusable assets, playbooks, user documentation, and implementation approaches that strengthen delivery excellence and collaborator enablement
  • Establish metrics to measure solution impact and value post-deployment and gather feedback for continuous improvement
  • Evaluate and foster the adoption of emerging AI-assisted development tools (e.g., code generation, review, testing automation) across the FDE team to accelerate delivery velocity and code quality
  • Know the latest digital trends in AI, software development, manufacturing, Industry 4.0 technologies, and relevant SaaS solutions
  • Ability to travel 25% to customer sites (will only require domestic US travel)

What you bring to the Team

  • Bachelor’s degree in Engineering, Computer Science, or related field required; MBA or relevant advanced degree; or equivalent combination of education and experience preferred
  • 8+ years of experience leading customer-facing software delivery, implementation, solution engineering, consulting, or technical delivery programs within enterprise SaaS, technology, or systems integrator environments
  • 2+ years of people leadership or technical leadership experience leading and developing engineers, architects, consultants, or technical delivery teams
  • 3-5 years of hands-on software development, solution engineering, integration development, or full-stack development experience, including technical leadership responsibilities
  • Experience with DevSecOps practices, secure software development lifecycle (SDLC) methodologies, and Agile/SAFe delivery frameworks
  • Fluency in AI technologies applied to software development lifecycle
  • Proficiency with project and delivery management tools such as Jira and Confluence, with a proven ability to lead cross-functional teams and multiple concurrent engagements
  • Experience in manufacturing operations, industrial engineering, or engineering services strongly preferred

The BIG payoff

We are a company who is going to make a difference in the industries and the communities in which we choose to operate.  Every employee of Re:Build will share ownership in the company and will share in the financial rewards of the success we achieve together, at all levels of the company!

We want to work with people that reflect the communities in which we operate

Re:Build Manufacturing is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, marital status, parental status, cultural background, organizational level, work styles, tenure and life experiences. Or for any other reason.

Re:Build is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need assistance or an accommodation due to a disability, you may contact us at accommodations.ta@ReBuildmanufacturing.com or you may call us at 617.909.6275.

Read the full description
Engineer Lead Data Engineer at Box

Designs and manages enterprise data platforms, lakehouse architectures, and data pipelines while leading teams on data governance, quality, and compliance.

Lead Posted 3 days ago RemoteFirstJobs Product
What this role involves

\*\*\* This is where your organization can create a consistent intro to all of your jobs, creating consistency in voice and messaging across all job posts

\*\*\* C’est ici que votre organisation peut créer une introduction cohérente à tous vos emplois, en créant une cohérence dans la voix et la messagerie dans tous les postes.

Marlabs, a global AI and Digital Solutions Consulting firm, delivers intelligent solutions across AI, data, analytics, and product engineering. Since 2000, we have partnered with some of the largest healthcare, life sciences, financial services, and government organizations worldwide. As we continue to expand our global footprint, we have an exciting opportunity for a highly skilled Lead Data Engineer to join our innovative and dynamic team.

Lead Data Engineer | About You

As a Lead Data Engineer, you will be responsible for designing, building, and managing the organization’s modern data platform, ensuring reliable, secure, and scalable data products that support analytics, reporting, and AI-driven business initiatives. You will lead the development of enterprise data pipelines, lakehouse architecture, and governance frameworks while partnering closely with AI/ML, platform engineering, and security teams. The ideal candidate combines deep expertise in data engineering, data modeling, cloud-based architectures, and data governance with a strong focus on reliability, observability, and regulatory compliance.

Lead Data Engineer | Day-to-Day

  • Design, build, and maintain scalable lakehouse architectures and enterprise data platforms that support analytics, reporting, and AI-driven solutions.
  • Develop and manage secure data ingestion frameworks, including CDC, batch, API, and file-based integrations from operational and transactional source systems.
  • Create and maintain data models, semantic layers, and governed metrics that enable consistent, trusted, and business-ready data consumption.
  • Implement data quality, reconciliation, observability, and monitoring processes to ensure reliable, recoverable, and high-performing data pipelines.
  • Partner with AI/ML, platform engineering, and security teams to deliver governed data products, support regulatory compliance requirements, and ensure proper data classification and access controls.
  • Establish and enforce data governance standards, lineage documentation, data contracts, quality thresholds, and operational procedures while supporting onboarding of new data sources and environments.

Lead Data Engineer | Skills & Experience

  • 8+ years of experience in Data Engineering, including end-to-end ownership of data ingestion, transformation, storage, and analytics delivery, with experience leading large-scale data initiatives.
  • Strong expertise in Python and SQL, including advanced data modeling, transformation frameworks, data quality management, and performance optimization.
  • Hands-on experience with modern lakehouse architectures and data platforms, including technologies such as Apache Iceberg, Delta Lake, Hudi, object storage, and cloud-native data solutions.
  • Proven experience building scalable data pipelines and CDC solutions, leveraging technologies such as Kafka, Debezium, Airflow, Dagster, dbt, and enterprise integration frameworks.
  • Strong understanding of data governance, lineage, security, and compliance practices, including data contracts, access controls, observability, audit readiness, and regulated industry environments.
  • Experience collaborating with AI/ML, analytics, platform engineering, and business teams to deliver trusted, governed, and scalable data products; financial services or banking industry experience is highly preferred.

\*\*\* Similar to the introduction that can precede all job descriptions, an outro can also be formatted for consistency on all posts

\*\*\* Semblable à l’introduction qui peut précéder toutes les descriptions de poste, une outro peut également être formatée pour la cohérence sur tous les messages

Read the full description
Engineer Staff Software Engineer at Fluxon

Staff-level engineer who writes production code, leads technical projects, mentors engineers, and shapes architectural decisions across multiple technology stacks.

Lead Remote Posted 3 days ago RemoteFirstJobs Product
What this role involves

Who we are

At Fluxon, we believe that how you build matters as much as what you build. We help businesses navigate their most important technology decisions with confidence, and take responsibility for seeing them through. Founded by ex-Googlers and startup veterans, we’re proud to partner with teams behind some of the most ambitious products, including Google, OpenAI, Anthropic, Walmart and Stripe.

Our work spans strategy, design, and engineering — often in complex, AI-driven environments — where clarity, speed and quality are the standard. We use AI intentionally, applying it only where it adds real value and expands what’s possible. Care shapes everything we do.

Inside Fluxon, you’ll find a global, remote-first team of experienced builders, who are curious, kind and serious about their craft. We’re building a place where people can take ownership, solve problems that matter and do work they’re proud to stand behind. If you want to do your best work alongside people who care as much as you do, you’ll feel at home here.

This role is fully remote, with candidates based in Buenos Aires, Argentina.

About the role

As a Staff Software Engineer at Fluxon, you’ll play a key role in shaping the technical direction of our engineering organization. This is a highly senior, hands-on leadership position where you’ll partner closely with company and engineering leadership to influence strategy, guide architectural decisions, and elevate our overall engineering practice. All Staff Engineers write production code, and everyone joins Fluxon as an individual contributor before stepping into project leadership or management.

You’ll be responsible for:

  • Guiding product delivery all the way to the user, leading projects, providing technical guidance, and building and iterating in a dynamic environment
  • Partnering directly with clients to understand their needs and achieve business goals
  • Defining product requirements, identifying appropriate system designs and planning development in partnership with our Product and Design teams
  • Helping drive a healthy and effective engineering culture within customer teams and inside Fluxon
  • Mentoring engineers across multiple teams, supporting their ongoing growth and strengthening team capabilities

You’ll work with a diversity of technologies, including:

  • Core Languages & Runtimes

    • Primary Languages: TypeScript/JavaScript, Python, Golang (Go), Java, C# (.NET), Kotlin, Swift, Rust
    • Additional: Ruby on Rails, Java, C# (.NET), Kotlin, Swift, Rust
  • Frameworks & Ecosystems

    • Front-End & UI: React, Next.js (Full-Stack), Angular, SwiftUI
    • Back-End & Server-Side: Spring Boot (Java/Kotlin), FastAPI (Python), Django (Python)
    • Mobile Development: Expo (React Native)
  • Cloud & Infrastructure

    • Platforms: Google Cloud Platform (GCP), Amazon Web Services (AWS), Microsoft Azure
    • Compute: GCP Compute Engine (VMs), AWS Fargate (Container Orchestration), Google Cloud Run (Serverless Containers), AWS Amplify
    • Storage: AWS S3 (Object Storage), Google Cloud Storage (GCS)
  • Data & Messaging Services

    • Streaming & Queuing: Apache Kafka (Event Streaming), AWS SQS (Simple Queue Service)
    • Data Warehouse & Analytics: Google BigQuery
    • Monitoring & Observability: GCP Cloud Monitoring Suite (CMS)
  • Data Stores

    • Relational (SQL): PostgreSQL, MariaDB
    • NoSQL/Document: Firestore (Firebase), Supabase, MongoDB
    • In-Memory Caching: Redis, Memcache
  • Advanced Technologies & Architecture

    • Artificial Intelligence/Machine Learning: AI/ML, Large Language Models (LLMs), Agentic AI, Natural Language Processing (NLP)
    • LLM Platforms: Google Gemini, OpenAI ChatGPT, Vertex AI (GCP), Anthropic Claude, Hugging Face (OSS Models)
    • Software Design: Architecture Redesign, Single Page Applications (SPA), Mobile Application Development
    • Emerging Tech: Blockchain/Crypto

Qualifications

  • 7+ years of industry experience in software development
  • Experience leading development through the full product lifecycle, including CI/CD, testing, release management, deployment, monitoring and incident response
  • Fluent in the design and implementation of scalable system architectures, data structures and algorithms, and effective development practices
  • Professional fluency in written and spoken English, with the ability to communicate technical concepts clearly to both international teammates and customers

What we offer

  • Remote-first, flexible work with a budget to set up a work space that works for you
  • Localized healthcare coverage to support you and your family’s wellbeing
  • Flexible paid time off with a minimum of 3 weeks per year (plus holidays)
  • “No internal meetings” Fridays, so you can focus on deep, uninterrupted work
  • A professional growth budget for learning that matters to you –  whether it’s developing your technical skills or learning a new language
  • A monthly wellness allowance to support your physical and mental health
  • Annual company offsites, where we gather in person to build meaningful connections
  • A competitive salary that reflects your expertise, impact and experience
  • Profit-sharing, so you can benefit directly from the value we build together
  • A paid sabbatical program, designed for rest, renewal and fresh perspective

We believe diverse teams perform better, and an inclusive environment is essential to building a successful organization. We welcome applicants from all backgrounds, experiences, and perspectives. We are an equal opportunity employer and are committed to providing accommodations throughout the hiring process.

This role uses AI-assisted tools to support initial screening. All assessments and decisions are made by a human reviewer.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead designs and manages GPU infrastructure, Kubernetes deployment, and managed inference platform architecture for Tether's AI compute services.

Lead Remote Posted 3 days ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical lead who designs and manages GPU infrastructure and Kubernetes-based compute platforms for AI inference and containerized workloads.

Lead Remote Posted 3 days ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure, Kubernetes deployment, and platform observability for a distributed compute and inference platform.

Lead Remote Posted 3 days ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Senior Staff Software Engineer, Benefits at Gusto

Senior Staff Software Engineer designs and builds full-stack platform capabilities for benefits applications, focusing on scalable systems and AI-driven interfaces.

Lead Posted 3 days ago RemoteFirstJobs Product
What this role involves

About Gusto

At Gusto, we’re on a mission to grow the small business economy. We handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on their craft and their customers. With teams in Denver, San Francisco, and New York, we support more than 500,000 small businesses nationwide and are building a workplace that reflects the people we serve.

All full-time employees receive competitive base pay, benefits, and equity (RSUs) — because everyone who helps build Gusto should share in its success. Offer amounts are determined by role, level, and location. Learn more about our Total Rewards philosophy.

AI is a fundamental part of how work gets done at Gusto. We expect all team members to actively engage with AI tools relevant to their role and grow their fluency as the technology evolves. AI experience requirements vary by role and will be assessed during the interview process.

About the Role

As a Staff Software Engineer on the Benefits Advisory team, you will be directly accountable for key architectural improvements to Gusto’s benefits platform. This is a full-stack role where you will design and build platform capabilities that enable customers to explore, apply for, and maintain their benefits within your product area.

Your work will focus on transforming our benefits opportunity, shopping and renewal flows, creating clear system boundaries that enable efficient reuse and increased scalability, allowing the ability to offer a wider variety of products more quickly. You will design services that continue to scale the organization, with an emphasis on enabling novel AI-driven interfaces to deliver a delightful benefits shopping experience.

You’ll operate at the intersection of marketing, sales, operations and engineering.  If you are passionate about building highly scalable systems that can reason, predict, and personalize to unlock Gusto Benefits’ next phase of sustainable growth, we’d love to have you join our team.

About the Team

The Benefits Advisory team is building the next generation of infrastructure that powers benefits applications, making it easier than ever to give customers access to the absolute best benefits for their needs.

Our mission is to build a robust system that informs our customers of the most relevant benefits options, making it easier than ever to confidently fulfill benefits that fit our customers current and future needs. We’re building a world-class technology platform designed to understand customer context in real-time, and continuously improve through data and feedback. This means rapid experimentation, fast learning, and strong cross-functional partnership.

We prioritize quality, observability, and uptime because these intelligent systems are fundamental to Gusto’s growth and brand. We partner closely with Marketing, Sales, and Operations to build and connect the AI-powered tools they use every day.

Here’s what you’ll do day-to-day:

  • Architect and evolve our customer-facing web platforms with a focus on scale and performance

Build innovative AI interfaces to best assist our customers

  • Design, build, and maintain shared services to enable rapid iteration on product offerings
  • Implement and optimize key workflows for a best-in-class customer experience.
  • Write high-quality, well-tested code across the stack, leveraging AI-powered tools to accelerate  development and improve reliability
  • Support, mentor, and up-level fellow engineers on the team, helping to establish best practices for AI-native design patterns.
  • Partner cross-functionally with Marketing, Sales, Operations and leadership to translate business needs into technical solutions that fit our needs now and in the future.
  • Influence the technical roadmap for your area of the Benefits Advisory platform, ensuring your work aligns with Gusto’s AI-native strategy and team goals

Here’s what we’re looking for:

  • 8+ years of experience building scalable full-stack web applications with expertise in both frontend (React, TypeScript) and backend (API development, data modeling, database schema design)
  • Extensive Experience building AI interfaces, iterating on prompts and maintaining high quality evals
  • Extensive Experience applying AI tools to accelerate full-stack development, improve code quality, and build intelligent user experiences. Familiar with AI-driven personalization patterns and willing to experiment with AI techniques to improve the benefits shopping experience
  • Deep experience in CI/CD and observability best practices.
  • Ability to act as a thought partner for both technical (Engineering, AI/Data) and business (Sales, Operations) teams.
  • A balance of pragmatic execution and long-term architectural thinking to build a scalable, intelligent platform.
  • Experience enabling productivity across large engineering organizations.
  • Nice to have:
    • Direct familiarity with React, Ruby on Rails
    • Experience with robust recommendation systems
    • AI architecture and scalability for reliable platform services

Compensation

Our cash compensation amount for this role is targeted at $191,000/yr to $225,000/yr in Denver & most remote locations, and $225,000/yr to $265,000/yr for San Francisco, Seattle & New York. Stock equity is additional. Final offer amounts are determined by multiple factors including candidate experience and expertise and may vary from the amounts listed above.

Gusto has physical office spaces in Denver, San Francisco, and New York City. Employees who are based in those locations will be expected to work from the office on designated days approximately 2-3 days per week (or more depending on role). The same office expectations apply to all Symmetry roles, Gusto’s subsidiary, whose physical office is in Scottsdale.

Note: The San Francisco office expectations encompass both the San Francisco and San Jose metro areas.

When approved to work from a location other than a Gusto office, a secure, reliable, and consistent internet connection is required. This includes non-office days for hybrid employees.

Our customers come from all walks of life and so do we. We hire great people from a wide variety of backgrounds, not just because it’s the right thing to do, but because it makes our company stronger. If you share our values and our enthusiasm for small businesses, you will find a home at Gusto.

Gusto is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristic. Gusto considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Gusto is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. We want to see our candidates perform to the best of their ability. If you require a medical or religious accommodation at any time throughout your candidate journey, please fill out this form and a member of our team will get in touch with you.

Gusto takes security and protection of your personal information very seriously. Please review our Fraudulent Activity Disclaimer.

Personal information collected and processed as part of your Gusto application will be subject to Gusto’s Applicant Privacy Notice.

Read the full description