The posting
Company Overview Docusign brings agreements to life. Over 1.5 million customers and more than a billion people in over 180 countries use Docusign solutions to accelerate the process of doing business and simplify people’s lives. With intelligent agreement management, Docusign unleashes business-critical data that is trapped inside of documents. Until now, these were disconnected from business systems of record, costing businesses time, money, and opportunity. Using Docusign’s Intelligent Agreement Management platform, companies can create, commit, and manage agreements with solutions created by the #1 company in e-signature and contract lifecycle management (CLM). What you'll do We're building the foundation that lets every engineering team at Docusign build, ship, and operate AI-native and agentic systems safely and at scale. This isn't a traditional infrastructure or DevOps role: it's an architecture and orchestration role for someone who has already designed and run production LLM and agentic systems at scale. As the AI Platform Lead, you own the architecture of our AI platform — the infrastructure, model gateway and routing layer, the agent orchestration frameworks, and the inference and GPU serving infrastructure that other teams build on top of. You set the technical direction, write the reference implementations and standards, and partner directly with product and platform engineering teams to get them adopted. You are deeply hands-on: you prototype, you write production code, and you use AI coding assistants and agentic tooling as a core part of how you work — but your primary output is the architecture and platform other engineers build against, not a single service. The successful candidate has real, hands-on experience building or operating LLM-powered and agentic services in production — not just experimenting with them. You're self-directed, comfortable owning ambiguous technical direction, and able to explain complex AI-infrastructure tradeoffs to both engineers and non-technical stakeholders. This position is an individual contributor role reporting to the Sr. Director of Cloud Services. Responsibility Own the end-to-end design, build, and integration of the AI/agentic platform that compliments human-operated execution with systems that route, remediate, and provision themselves — model gateway, agent orchestration, inference serving — and the standards other teams build against. You write the code; you don't hand it off Architect, build, and operate a multi-provider LLM gateway and routing layer (e.g., LiteLLM, Portkey, Bedrock, or equivalent) so requests reach the right model automatically — secure, cost-aware, and reliable — without a person routing them by hand Design and build the orchestration layer and frameworks (e.g., LangGraph, CrewAI, or a custom runtime) that let agents carry out work autonomously, and define the rules — capability schemas, guardrails, policy enforcement — those agents operate under Architect and operate scalable model-serving and GPU infrastructure (e.g., vLLM, KServe/Triton, Kubernetes-based inference serving, autoscaling, capacity planning) that balances latency, throughput, and cost Replace manual provisioning and ticket-driven execution with automated, self-service systems — teams get what they need and common failures resolve themselves, without a person in the loop Build golden paths for fine-tuning, model/agent registries, evaluation, and safe deployment — including CI/CD for models, prompts, and agents with canary, A/B, and shadow rollouts Define and evangelize architecture patterns, security guardrails, and reusable accelerators/reference architectures that reduce time-to-value for teams building on the platform Partner closely with product and platform engineering teams to drive adoption of the AI platform, unblock their AI/agentic use cases, and translate platform capabilities into their roadmaps Implement security controls and Zero Trust practices — including SBOMs, image signing, and policy-as-code — partnering with InfoSec to maintain a strong security posture across model access, data handling, and agent permissions Establish observability for AI infrastructure — token spend, model latency, GPU utilization, agent behavior — and continuously verify that what you've automated performs as intended Job Designation Hybrid: Employee divides their time between in-office and remote work. Access to an office location is required. (Frequency: Minimum 2 days per week; may vary by team but will be weekly in-office expectation) Positions at Docusign are assigned a job designation of either In Office, Hybrid or Remote and are specific to the role/job. Preferred job designations are not guaranteed when changing positions within Docusign. Docusign reserves the right to change a position's job designation depending on business needs and as permitted by local law. What you bring Basic Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience 12+ years of experience in Cloud, Platform, or Infrastructure Engineering with a Bachelor's degree or 8+ years of experience with a Master's degree Experience architecting and operating LLM-powered or agentic systems in production (not just prototypes or POCs) Experience replacing manual, ticket-driven operations with automated or self-service systems — building the systems that act on what they observe, not just dashboards and monitors Experience designing and operating an LLM gateway / model routing layer across multiple providers (e.g., OpenAI, Anthropic, Bedrock, Azure OpenAI) Experience building and operating agent orchestration frameworks and agentic workflows (e.g., LangGraph, CrewAI, AutoGen, or a custom orchestration runtime) in production Experience with GPU/inference-serving infrastructure and Kubernetes at scale (e.g., vLLM, KServe, Triton, model autoscaling, capacity planning) Experience writing, reviewing, and testing production-quality code (Python and/or Go) — this is a build role, not a slide-deck role Experience with cloud platforms (AWS, Azure, or GCP) and Infrastructure-as-Code (Terraform or Ansible) Experience setting technical direction across teams and driving adoption of a shared platform, with strong written and verbal communication Preferred Experience with RAG and vector database infrastructure (embedding pipelines, retrieval services) in production Experience running LLMOps practices — model evaluation, prompt/version management, guardrails, and agent monitoring Experience with containerization and microservices architecture (Docker, Kubernetes) beyond AI workloads Understanding of network architecture and security best practices (VPNs, firewalls, load balancing) in cloud environments Experience presenting architecture and roadmap to senior technical leadership and driving org-wide adoption Wage Transparency Pay for this position is based on a number of factors including geographic location and may vary depending on job-related knowledge, skills, and experience. Based on applicable legislation, the below details pay ranges in the following locations: California: $164,700.00 - $266,000.00 annual base salary This role is also eligible for the following: Bonus: Sales personnel are eligible for variable incentive pay dependent on their achievement of pre-established sales goals. Non-Sales roles are eligible for a company bonus plan, which is calculated as a percentage of eligible wages and dependent on company performance. Stock: This role is eligible to receive Restricted Stock Units (RSUs). Global benefits provide options for the following: Paid Time Off: earned time off, as well as paid company holidays based on region Paid Parental Leave: take up to six months off with your child after birth, adoption or foster care placement Full Health Benefits Plans: options for 100% employer paid and minimum employee contribution health plans from day one of employment Retirement Plans: select retirement and pension programs with potential for employer contributions Learning and Development: options for coaching, online courses and education reimbursements Compassionate Care Leave: paid time off following the loss of a loved one and other life-changing events Work Authorization Notice: Please note that we do not provide visa sponsorship or immigration support for this position. Applicants must already be authorized to work in the United States on a full-time, permanent basis without the need for current or future sponsorship. Life at Docusign Working here Docusign is committed to building trust and making the world more agreeable for our employees, customers and the communities in which we live and work. You can count on us to listen, be honest, and try our best to do what’s right, every day. At Docusign, everything is equal. We each have a responsibility to ensure every team member has an equal opportunity to succeed, to be heard, to exchange ideas openly, to build lasting relationships, and to do the work of their life. Best of all, you will be able to feel deep pride in the work you do, because your contribution helps us make the world better than we found it. And for that, you’ll be loved by us, our customers, and the world in which we live. Accommodation Docusign is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need such an accommodation, or a religious accommodation, during the application process, please contact us at [email protected]. If you experience any issues, concerns, or technical difficulties during the application process please get in touch with our Talent organization at [email protected] for assistance. Applicant and Candidate Privacy Notice States Not Eligible for Employment This position is not eligible for employment in the following states: Alaska, Hawaii, Maine, Mississippi, North Dakota, South Dakota, Vermont, West Virginia and Wyoming. Equal Opportunity Employer It's important to us that we build a talented team that is as diverse as our customers and where all employees feel a deep sense of belonging and thrive. We encourage great talent who bring a range of perspectives to apply for our open positions. Docusign is an Equal Opportunity Employer and makes hiring decisions based on experience, skill, aptitude and a can-do approach. We will not discriminate based on race, ethnicity, color, age, sex, religion, national origin, ancestry, pregnancy, sexual orientation, gender identity, gender expression, genetic information, physical or mental disability, registered domestic partner status, caregiver status, marital status, veteran or military status, or any other legally protected category. EEO Know Your Rights poster #LI-Hybrid
We're building the foundation that lets every engineering team at Docusign build, ship, and operate AI-native and agentic systems safely and at scale. This isn't a traditional infrastructure or DevOps role: it's an architecture and orchestration role for someone who has already designed and run production LLM and agentic systems at scale. As the AI Platform Lead, you own the architecture of our AI platform — the infrastructure, model gateway and routing layer, the agent orchestration frameworks, and the inference and GPU serving infrastructure that other teams build on top of. You set the technical direction, write the reference implementations and standards, and partner directly with product and platform engineering teams to get them adopted. You are deeply hands-on: you prototype, you write production code, and you use AI coding assistants and agentic tooling as a core part of how you work — but your primary output is the architecture and platform other engineers build against, not a single service. The successful candidate has real, hands-on experience building or operating LLM-powered and agentic services in production — not just experimenting with them. You're self-directed, comfortable owning ambiguous technical direction, and able to explain complex AI-infrastructure tradeoffs to both engineers and non-technical stakeholders. This position is an individual contributor role reporting to the Sr. Director of Cloud Services. Responsibility Own the end-to-end design, build, and integration of the AI/agentic platform that compliments human-operated execution with systems that route, remediate, and provision themselves — model gateway, agent orchestration, inference serving — and the standards other teams build against. You write the code; you don't hand it off Architect, build, and operate a multi-provider LLM gateway and routing layer (e.g., LiteLLM, Portkey, Bedrock, or equivalent) so requests reach the right model automatically — secure, cost-aware, and reliable — without a person routing them by hand Design and build the orchestration layer and frameworks (e.g., LangGraph, CrewAI, or a custom runtime) that let agents carry out work autonomously, and define the rules — capability schemas, guardrails, policy enforcement — those agents operate under Architect and operate scalable model-serving and GPU infrastructure (e.g., vLLM, KServe/Triton, Kubernetes-based inference serving, autoscaling, capacity planning) that balances latency, throughput, and cost Replace manual provisioning and ticket-driven execution with automated, self-service systems — teams get what they need and common failures resolve themselves, without a person in the loop Build golden paths for fine-tuning, model/agent registries, evaluation, and safe deployment — including CI/CD for models, prompts, and agents with canary, A/B, and shadow rollouts Define and evangelize architecture patterns, security guardrails, and reusable accelerators/reference architectures that reduce time-to-value for teams building on the platform Partner closely with product and platform engineering teams to drive adoption of the AI platform, unblock their AI/agentic use cases, and translate platform capabilities into their roadmaps Implement security controls and Zero Trust practices — including SBOMs, image signing, and policy-as-code — partnering with InfoSec to maintain a strong security posture across model access, data handling, and agent permissions Establish observability for AI infrastructure — token spend, model latency, GPU utilization, agent behavior — and continuously verify that what you've automated performs as intended
Basic Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience 12+ years of experience in Cloud, Platform, or Infrastructure Engineering with a Bachelor's degree or 8+ years of experience with a Master's degree Experience architecting and operating LLM-powered or agentic systems in production (not just prototypes or POCs) Experience replacing manual, ticket-driven operations with automated or self-service systems — building the systems that act on what they observe, not just dashboards and monitors Experience designing and operating an LLM gateway / model routing layer across multiple providers (e.g., OpenAI, Anthropic, Bedrock, Azure OpenAI) Experience building and operating agent orchestration frameworks and agentic workflows (e.g., LangGraph, CrewAI, AutoGen, or a custom orchestration runtime) in production Experience with GPU/inference-serving infrastructure and Kubernetes at scale (e.g., vLLM, KServe, Triton, model autoscaling, capacity planning) Experience writing, reviewing, and testing production-quality code (Python and/or Go) — this is a build role, not a slide-deck role Experience with cloud platforms (AWS, Azure, or GCP) and Infrastructure-as-Code (Terraform or Ansible) Experience setting technical direction across teams and driving adoption of a shared platform, with strong written and verbal communication Preferred Experience with RAG and vector database infrastructure (embedding pipelines, retrieval services) in production Experience running LLMOps practices — model evaluation, prompt/version management, guardrails, and agent monitoring Experience with containerization and microservices architecture (Docker, Kubernetes) beyond AI workloads Understanding of network architecture and security best practices (VPNs, firewalls, load balancing) in cloud environments Experience presenting architecture and roadmap to senior technical leadership and driving org-wide adoption



