Staff Engineer responsible for technical operations, governance, and incident response for enterprise workplace AI systems including platform configuration, integrations, observability, and production support.
Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and synthetic reality technologies. With offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, Shield AI’s technology actively supports operations worldwide. For more information, visit www.shield.ai. Follow Shield AI on LinkedIn, X, Instagram, and YouTube.
This is a deeply technical individual contributor role reporting to the Head of AI Operations & Governance (Enterprise AI). This person will own significant portions of the technical operating load for workplace AI: platform configuration, connector and integration management, observability wiring, secrets and access hygiene, model/prompt lifecycle mechanics, and hands-on changes in production.
At the same time, the role will help translate production realities into governance artifacts, risk assessments, status updates, and training/enablement support — giving the Head a force multiplier who can operate at both the technical and “softer” layers of AI operations and governance.
This role will also serve as a hands-on technical responder for workplace AI incidents and production issues. The Staff Engineer is expected to investigate failures, troubleshoot across platforms, integrations, prompts, configurations, and access paths, implement mitigations or fixes where appropriate, and help restore service quickly while documenting root cause, lessons learned, and prevention steps.
Own day-to-day configuration of AI platforms and orchestration tools (models, routes, guardrails, tenants, policies, role mappings, prompt libraries, etc.), under the direction of the Head.
Design, configure, and maintain connectors and extensions into SaaS systems, data sources, and workflow tools; ensure connectivity is reliable, secure, and aligned with access policies.
Set up and maintain logging, metrics, and alerts for AI workflows and tools; make sure key signals (latency, errors, usage, drift indicators) are captured and visible to the team.
Implement secure storage and rotation for API keys, tokens, and credentials; maintain access control configurations and partner with Security/IT on reviews and remediation.
Implement model swaps, policy updates, prompt changes, version upgrades, and rollout plans based on decisions made by the Head; maintain detailed change records and rollback paths.
Run experiments and benchmarks on models, tools, and configurations; collect and summarize technical performance data to inform governance and roadmap decisions.
Help the Head interpret logs, metrics, and incidents into clear risk, reliability, and compliance narratives that can be shared with Security, Legal, HR, and business sponsors.
Co-author technical sections of operational playbooks, runbooks, and training materials; occasionally participate in training or office hours to help users understand capabilities and guardrails.
Act as a technical point of contact in incidents: triage, investigate, propose mitigations, execute fixes, and document learnings; communicate clearly with non-technical stakeholders when needed.
#LI-KE1
#LC
Pay within range listed + Bonus + Benefits + Equity
Pay within range listed above + temporary benefits package (applicable after 60 days of employment)
Salary compensation is influenced by a wide array of factors including but not limited to skill set, level of experience, licenses and certifications, and specific work location. All offers are contingent on a cleared background and possible reference check. Military fellows and part-time employees are not eli
Sourced from the a16z Speedrun talent network — apply on Speedrun.