The posting
ABOUT MARIANA MINERALS
Mariana Minerals is a software-first, vertically integrated minerals company on a mission to supply the critical minerals powering modern energy, AI, and defense technologies. We’re reimagining the minerals supply chain by combining deep industry expertise with advanced software, automation, and data-driven decision-making.
THE ROLE
Mariana Minerals is looking for a Sr. or Staff Software Engineer, Infrastructure to join the team that owns and collaborates on the platform every other engineer here builds on.
Our software team builds MarianaOS, the operating system for how Mariana engineers, constructs, and operates its mines and mineral processing facilities. It runs in AWS and on edge clusters at our sites, including our copper site in Moab, Utah and our lithium facility in East Texas. On any given day, it serves construction dashboards, running GPU inference on site camera feeds, training models for a lithium extraction circuit, and moving plant floor data into a warehouse.
Your team owns the platform underneath all of that: clusters, deployment, CI/CD, observability, site/edge footprint, identity, secrets, and cloud costs. You will also collaborate with machine learning engineers and data engineers on the ML compute platform and data pipeline infrastructure.
The part of this role we care most about is what you build for other people. We do not want an approval queue. When an engineer needs a database, a bucket, a dashboard, or an alert, they should get it from a documented path they can run themselves, and your job is to make that path exist and make it the easiest option. You will be measured on how much infrastructure work happens without you.
We move quickly. Our engineers ship, get feedback from the people running our plants and job sites, and iterate. Everything you build should shorten that loop.
Our stack is TypeScript and React on the front end, Node and Python services, Postgres, and AWS. The platform is EKS with Karpenter, ArgoCD, Helm, Pulumi, Prometheus and Grafana, Loki, Prefect, Snowflake, and a growing amount of LLM and agent tooling in production.
WHAT YOU'LL DO
Own our Kubernetes footprint end to end: the EKS cluster, the on-premises site clusters, cluster upgrades, capacity and autoscaling, ingress and DNS, and the GitOps delivery path that turns a merged commit into a running service.
Build self-service infrastructure for other engineers, so that provisioning a service, a database, a preview environment, or a scheduled job is a documented path they run themselves so the team can deliver quickly.
Make observability a default. Every deployed service should emit metrics, ship logs, trace requests, and carry meaningful alerts without its author's writing bespoke plumbing.
Extend the platform to our sites, running on-premises and edge clusters that host GPU inference and plant data collection over constrained links.
Own cloud cost as an engineering problem, with the visibility, budgets, and right-sizing to keep a growing platform from growing its bill proportionally.
Own identity, access, and secrets management, so that scoped access is provisioned and revoked as a matter of course.
Respond to infrastructure incidents and then remove the class of failure that caused them, including writing the runbook so the fix scales.
WHAT YOU'LL BRING
8-12+ years building and operating production infrastructure, with real depth in Kubernetes: you have run and upgraded clusters yourself and debugged multiple layers below the application.
Strong software engineering fundamentals. You write code you would be comfortable having reviewed by a product engineer, and you reach for a well-designed tool or abstraction before a runbook.
Hands-on experience with infrastructure as code (Pulumi, Terraform, or similar) and GitOps-style continuous delivery.
A track record of building platforms other engineers adopted willingly. We will ask you for specifics: what you built, who used it, and how you know it helped.
Practical observability experience with Prometheus, Grafana, and a log aggregation stack, and getting an organization to instrument its services.
Comfort operating across a broad surface with incomplete context.
Fluency in Python or TypeScript, and enough familiarity with AWS to make sound cost and architecture tradeoffs.
Experience supporting ML or data-intensive workloads (GPU scheduling, distributed training, pipeline orchestration) is a strong plus.
Experience with on-premise or edge infrastructure, industrial networks, or constrained-connectivity environments is a strong plus.
Startup experience in a rapidly scaling company is helpful but not required.
Domain experience in manufacturing, industrial automation, mining, or capital projects is beneficial but not required.
HOW WE'LL INTERVIEW
We want to see you design and build platform tooling, so the loop is weighted toward that.
Expect a conversation about a platform you have built and why you made the calls you made, a working session where you design a self-service capability for an engineering org and we push on the tradeoffs, and a hands-on exercise with infrastructure code.
OUR CULTURE IS BUILT ON FOUR PRINCIPLES:
Everyone Gets Home Safe. We never put speed or cost ahead of people.
Extreme Ownership. We take full responsibility for outcomes, relentlessly driving toward solutions.
Engineer Out Requirements, then Automate. We simplify, optimize, and then automate for scale.
Share Your Legos. We collaborate openly, share knowledge, and empower each other to build bigger, better solutions.
Join us as we build the future of responsible mineral sourcing and supply!



