Principal AI Platform Engineer
Iconma Portal
4 days ago
Remote
$75.59 - $80.59 USD hourly
Web Development
Our client, a Digital Infrastructure company, is looking for a Principal AI Platform Engineer for their Remote location.
Responsibilities:
Requirements:
Why Should You Apply?
Responsibilities:
- The ISD team runs its delivery work through a shared library of Claude toolkits: pricing engines, deployment pipelines, dashboards, a reference architecture catalog, and an Obsidian knowledge base, all wired to Smartsheet and SharePoint and driven by six named Claude agents.
- It works. It is also a 692 MB folder on SharePoint with no version control, where the code, the customer data and the published outputs sit in the same tree, and where every routine task loads a large part of the library into a Claude session before it does anything. Four people can drive it. The rest of the team cannot.
- We need someone to own that platform and make it usable by the whole team, not just the people who can read the code. Four jobs.
- Keep it running and extend it.
- SharePoint and Smartsheet integrations, the dashboards and reports the team publishes, the shared Python library underneath. This is the standing half of the role and it does not stop while the rest happens.
- Take the toolkits off the context window.
- Every toolkit today is markdown instructions and Python that a Claude session reads before it does any work, so a routine pricing run or a data refresh costs tokens in proportion to the size of the library rather than the size of the job. The work is to measure where the tokens actually go, then move the routine paths behind a callable surface so the model invokes a typed tool instead of reading the library. An MCP server, a packaged plugin with progressive disclosure, a headless Agent SDK service, and a small internal web app are all on the table. We have not picked one, and a candidate who arrives having picked one without measuring is answering a question we have not asked. The target is a stated reduction in cost and time per run against a baseline you establish in the first month.
- Run it in GitHub, with the data kept out.
- The library has no version control today. Own the migration, the CI, and packaged releases of the skills and agents. The harder half is the boundary: the code goes to GitHub, and the customer pricing data, the deal content, the knowledge base and the published outputs stay in Client systems and never reach a repo. You design that split, set up the scanning that enforces it, and solve distribution, because a consumer on a fresh Mac has to be able to install a release and point it at SharePoint data without asking an engineer.
- Open the knowledge base to the team.
- The ISD Brain is 926 wiki pages built from 875 source documents and 377 meeting summaries, with citation rules, source confidence tiers and restricted pages that must not leave the vault. One person ingests into it today and one person queries it well. The job is to make it a retrieval surface the whole team uses and contributes to, which means multi-user writes on a SharePoint-synced vault, access control that holds for restricted pages, grounded answers that cite the page they came from, and a system that says "I don't know" instead of filling the gap.
- This is hands-on build work. The deliverable is working code and documentation someone can follow without asking
Requirements:
- Toolkits: 9 registered toolkits plus a shared _platform library, ~200 Python files, ~62,000 lines, 33 skills, 6 named agents
- Version control None today. A 692 MB SharePoint folder synced to each person's Mac
- Knowledge base ISD Brain, an Obsidian vault of 926 markdown pages, 875 raw sources, 377 meeting summaries, 836 MB
- Systems of record Smartsheet for REQ tracking, price sheets and the backlog; SharePoint for the library and published outputs
- Access model Four roles already defined: owner, maintainer, operator, consumer. Writes to a system of record are gated on a human yes
- Users A small team. Most are not engineers and will never read the code
- Claude Code, at depth.
- You have built with it, not just used it. Skills, sub-agents, CLAUDE.md instruction files, hooks, permissions, MCP connectors, and the token behaviour underneath all of it. This is the primary tool of the role and the first thing we will ask you to show us.
- GitHub and CI.
- You have migrated an existing body of work into version control and kept its history and attribution intact. You can design a repo structure, a branching model and a release process that a small non-engineering team will actually follow, and you know how to keep data and secrets out of a repo by construction rather than by reminder.
- Python.
- Confident writing and testing pipelines with the standard library. You treat a new dependency as a cost to justify, not a default, because anyone on the team has to be able to run this from a fresh machine.
- Experience building with AI tools.
- Real delivery with agentic systems and RAG. You know how to draw boundaries between agents, ground answers in cited sources, and make a system say "I don't know" instead of guessing. You can show what a system cost to run before and after you worked on it.
- Handling of confidential material.
- The toolkits carry customer pricing, deal content, partner terms and named-account detail, and the knowledge base carries pages that are scoped to a working team. You treat a folder boundary as a security control, not a filing convention, and you assume nothing about access that was not granted to you.
- Respect for provenance.
- Everything this team publishes carries a number that someone will act on. You treat an unsourced or undated figure as a defect.
- Jira and Confluence.
- You manage a backlog and write design docs and runbooks that other people use.
- Translating technical work into business value.
- The people who use and fund this are not engineers. You can say what a pipeline or an agent is worth in terms they care about: the deal it prices faster, the deployment it de-risks, the hours it gives back. Work that cannot be explained that way is work worth questioning.
- Requirements arrive here as a request for a dashboard, a script, or a column. You find the decision behind the request before you build, and you say something when the thing asked for will not solve the problem described.
- Most of this work sits between someone else's domain knowledge and the code. You draw it out, work in the open, share a rough version early rather than a finished one late, and take a correction without defending the first attempt.
- A large share of the output is prose read by people with no engineering background. It has to be clear and concise.
- Obsidian or a comparable markdown knowledge base at team scale
- SharePoint, and Microsoft 365 connectors
- Smartsheet, Confluence
- API and MCP Connectivity
- GCP or AWS
- Shared code affects everyone, so you ask before changing it. Every publish and every write to a system of record is gated on a human yes. Credentials live in the OS keychain and never in a file. Some toolkits and some knowledge-base pages are scoped or restricted, and access is granted, not assumed
Why Should You Apply?
- Health Benefits
- Referral Program
- Excellent growth and advancement opportunities