The AI ops role is the job your company created without meaning to. Somewhere between the first agent going live and the third one quietly breaking, somebody started owning the prompts, the integrations, the API spend, the incident at four on a Friday, and the vendor who deprecated a model with two weeks of notice. That person usually has a different title on their contract, no budget line, and no plan. This guide gives the job a name, a band, and a first ninety days.
The timing is not academic. In a survey of 942 US small-business owners fielded on April 7 to 9, 2026 and published by Bluevine in July, 74% said they were already using or actively testing AI tools, while only 22% were completely confident AI could handle low-level tasks without human supervision. Read those two numbers together and the job description writes itself: most companies are running the thing, and almost nobody trusts it to run alone. Supervision is a person, and a person is a line in a budget.
What follows is the practical version for a company of twenty to two hundred people. What the job actually covers, what the six competing titles mean and which one to put in a posting, what the published pay data does and does not tell you, a decision table for hiring versus promoting internally versus going fractional, a first ninety days, and the failure mode we see more than any other — handing the whole thing to one already-busy engineer as an unfunded side-quest.
- 01The job exists before the title does.At least six labels compete for the same scope at the time of writing — AI Agent Manager, AI Ops Manager, Agent Operator, Agent Supervisor, Agent QA Lead, AI Champion — with no market consensus. Pick the title that matches where the role reports, not the one that sounds most senior.
- 02Published pay spans about 3.3 times, floor to ceiling.A $75,000 job-board band floor against a $247,000 posted base at an Alphabet subsidiary is not one market pricing one job. It is at least three different jobs wearing one name, and the spread tracks who the role reports to.
- 03Fractional is the fast option, not the cheap one.The cheapest individual fractional retainer in our source set annualises to $84,000 — $9,000 more than the floor of the cheapest full-time band before employer costs. Buy it for speed and senior judgement, not to save money.
- 04Domain knowledge beats AI credentials.Both priced live postings we found ask for years of business, data or consulting experience rather than an AI-specific degree, and one of the two also asks for hands-on generative AI prototyping. The person who already knows the workflow learns the tooling faster than the reverse.
- 05The dominant failure is an unfunded side-quest.Not a bad hire. One busy engineer given the work with no scope, no budget and no ninety-day plan, and a pilot approved without the production budget behind it. Decide who owns the thing after the pilot before the pilot starts.
01 — The SituationSomebody at your company already has this job.
The work arrives before the role does. A marketing manager wires up a drafting agent inside a tool the company already pays for. An engineer builds a triage bot over the ticket queue. Someone in finance points a model at a spreadsheet. Each one works well enough that nobody escalates it, and each one quietly acquires a dependency: a credential, a data source, a colleague who now assumes the output will be there on Monday.
Then a model version changes, or a connector breaks, or the monthly API bill triples, and the question surfaces for the first time: who owns this? The honest answer at most companies is that somebody does — informally, in the margins of another job, without a budget or a mandate. That informal owner is the AI ops role in its natural habitat, and formalising it is usually cheaper than replacing them after they leave for a posting that names the work.
The adoption data explains why this stopped being an enterprise-only problem. The Bluevine survey covered US owners and financial decision-makers at companies with two to 249 employees and $50,000 to $5 million in revenue, fielded by Centiment across three days in April 2026, with a published margin of error of about three percentage points at 95% confidence. That is a small-business sample, not a Fortune 500 one, and the picture it paints is of broad adoption sitting on top of thin supervision.
SMB AI adoption, and the supervision gap underneath it
Source: Bluevine, 2026 Small Business AI Trends Report — Centiment survey of 942 US small-business owners, fielded April 7–9, 2026, published July 15, 2026Note that the bars do not share a denominator, which is exactly the point. Adoption and barriers are measured across every respondent; return and time saved are measured only among the owners already using AI. So the 52% return figure is the optimistic cut — it excludes everyone who has not started. Among the people furthest along, roughly half can point to a measurable result and roughly a quarter cannot yet. The gap between those two groups is almost never about model quality. It is about whether anyone owns the deployment after the demo.
The most-cited use-case areas track the same story: data analysis and business insights at 39%, marketing and sales at 37%, automating operations at 29%, customer service at 27%, financial management at 23%. Those are five different departments, five different data sources, and five different definitions of an acceptable error. One person is going to have to hold all of them in their head. For the broader question of what to automate first, the sequencing playbook this hire executes covers the order of operations; this post is about who does the executing.
02 — NamingSix titles, one job, and one term that will sink your posting.
There is no settled name for this job. The same scope shows up as AI Agent Manager, AI Ops Manager, Agent Operator, Agent Owner, Agent Supervisor, Agent QA Lead and AI Champion, and the two priced live postings we found both used an eighth label — AI Enablement Lead. One industry analysis we leaned on for the taxonomy describes AI Agent Manager as the current front-runner, which is a reasonable read of a market that has not actually voted yet.
For an owner writing a job ad, the naming problem is not cosmetic. Applicant tracking systems, job-board alerts and recruiter searches all sort on the literal string. Choose a label nobody searches for and your posting is invisible; choose the wrong label and it lands in a completely different market. The table below is our decoder, built from the titles that actually appeared in the research rather than from a taxonomy anyone has ratified.
| Title | What it usually means | Typical seniority | Put it in your ad? |
|---|---|---|---|
| Front-runner labels | |||
| AI Agent Manager | The internal owner of agents in production: prompts, integrations, evals, escalation, model upgrades. | Manager to senior manager | Yes. Described as the front-runner in the analysis we reviewed, and the closest match to the full scope. |
| AI Ops Manager | The same scope in operations language — reliability, process, throughput, cost per workflow. | Manager | Yes, when the role reports into operations rather than IT. Attracts process people over platform people. |
| AI Enablement Lead | Adoption-first framing: workflow assessment, governance and privacy alignment, training, internal advocacy. | Lead to senior lead | Yes. Both of the dated, priced postings we located used this exact title, which makes it the most searchable of the three. |
| Narrower functional labels | |||
| Agent Operator | Day-to-day running of live agents. Popularised by public commentary rather than by job boards. | Individual contributor | In the body copy, not the headline. Recognisable to people who follow the space, invisible to everyone else. |
| Agent Supervisor | Monitoring, exception handling and quality checks on agent output. Narrower than ownership. | Individual contributor | Only if that genuinely is the job. It prices at the bottom of the reported bands for a reason. |
| Agent QA Lead | Evals, test sets and regression checks, especially after a model version changes underneath you. | Senior IC to lead | Yes, as a second hire — once one person can no longer both build the agents and honestly check them. |
| AI Champion | An internal advocate, almost always part-time and almost always unfunded. | Not a level | No. It is a behaviour, not a job. Titling it this way is how the role becomes a side-quest. |
| The label to avoid | |||
| AI Operator | On general job boards this is a labelling, monitoring and prompt-running gig, reported at roughly $19 to $25 an hour. | Hourly | No. Screening tools sort on the literal string, and this one sorts your posting into the wrong pile entirely. |
Our read on where this goes: the titles consolidate within a year or two, and the tell will be job boards adding a filter for it. Until then, the pragmatic move is to lead with the label that matches your reporting line, name two or three of the alternates inside the ad body so search picks them up, and stop worrying about it. Candidates who understand the work are already used to the naming chaos — several of them are doing this job today under a title like IT Manager or Operations Lead.
03 — The ScopeEight duties that cluster into four halves of one job.
The most useful public description of the scope comes from Box chief executive Aaron Levie, who published an eight-part specification for an internal agent operator: map new workflows around agents, implement the systems that deploy them, keep agent context current, wire internal systems into them, build evals, decide where humans stay in the loop, manage model-upgrade drift, and run change management. We are working from a secondary analysis of that specification rather than an independently audited standard, so treat the eight items as one well-placed operator’s framing rather than an industry consensus. As framings go it holds up, because every item on it maps to something that breaks in real deployments.
Grouped into pairs, the shape of the job becomes obvious — and it is not the shape most companies expect.
Map and deploy
Redesign the workflow around what an agent can actually do, then stand up the systems that let agents run inside it. This is the half that looks like operations rather than engineering, and it is where most of the value is decided before a single prompt is written.
Context and plumbing
Agents decay when the documents, records and policies behind them go stale. Somebody has to keep that context fresh and connect the agent to the systems that hold it. Unglamorous, continuous, and the first thing to lapse when the owner is part-time.
Evals and the human line
Write the tests that say whether an agent is still good, and draw the line where a human has to approve. Both are judgement calls about the business and its tolerance for error, not calls about the model. This is where the risk actually lives.
Drift and change
Model versions move underneath you and people have to be brought along. The second is the harder job by some distance — an agent that works and that nobody trusts is indistinguishable, operationally, from an agent that does not work.
Most of those eight duties are business judgement rather than technical implementation. That is the strongest argument for hiring from the function being automated rather than from an AI background, and the postings back it up: the Alphabet-subsidiary listing asks for eight to ten years of data, machine-learning or business-intelligence experience layered with hands-on generative AI prototyping, and the staffing-firm listing asks for technology, consulting, solutions architecture or technical program management experience. Neither asks for an AI-specific degree. The closest existing analogy in the research is a Salesforce administrator: someone who does not build the platform but configures it, keeps its data honest, trains the people using it, and decides what to switch on next.
The change-management duty deserves its own budget line rather than a bullet. If you want the full version of that half of the job, the change-management playbook for AI adoption is the companion piece. And if you are hiring engineers alongside this role, note that they are different searches with different screens — hiring for AI-adjacent technical roles is a genuinely separate exercise from hiring the person who runs what those engineers ship.
“We haven’t removed humans from the loop, we’ve just changed where they enter the loop.”— Aaron Levie, CEO of Box, on the 20VC podcast, April 2026
That line is the job description compressed into one sentence. Deciding where the human enters is not a setting in a dashboard. It is a per-workflow decision about what an acceptable error costs, made by someone who understands both the workflow and the failure modes of the model running it. Levie also predicted somewhere between 500,000 and a million jobs of this kind would be created as companies internalise the work — a founder’s forecast rather than a labour statistic, and worth reading as a directional claim about where the work lands rather than as a number to plan against.
04 — The MoneyA 3.3× spread means these are not the same job.
Published pay for this role is all over the place, and the spread is informative rather than noise. The table below normalises every figure we found onto the same two axes — per month and per year — so that a salary band, a live posting and a consulting retainer can finally be read against each other. Every derived cell is arithmetic on the published figure in the same row: annual divided by twelve, or monthly multiplied by twelve.
Read the source labels carefully, because they are not equivalent evidence. Posted base means a real advertisement with a number in it. Job-board reported means aggregated or self-reported data of unknown depth. Rate card means a consultancy publishing prices about its own market, which is the weakest surface of the three.
| Route | As published (and on what surface) | Per month | Per year |
|---|---|---|---|
| Employ — base salary only, before employer costs | |||
| Agent Supervisor band | $75,000–$120,000/yr · job-board reported band, attributed analysis, May 2026 | $6,250–$10,000 | $75,000–$120,000 |
| AI Ops Manager band | $110,000–$160,000/yr · job-board reported band, attributed analysis, May 2026 | $9,167–$13,333 | $110,000–$160,000 |
| AI Enablement Lead · staffing firm for an ed-tech client, Raleigh NC | $130,000/yr · posted base on a live listing dated August 4, 2026 | $10,833 | $130,000 |
| AI Enablement Lead · Waymo, Mountain View | $200,000–$247,000/yr plus bonus and equity · posted base, listing dated July 2026 | $16,667–$20,583 | $200,000–$247,000 |
| Glassdoor “AI Operations Manager” | Two search cuts averaging $122,796 and $161,870/yr · each based on roughly one self-reported submission, so illustrative only | $10,233–$13,489 | $122,796–$161,870 |
| Contract — consulting-market rate cards, vendor-adjacent | |||
| Fractional AI architect · 4–8 hrs/week | $7,000–$15,000/mo · consultancy rate card published June 2, 2026, derived from a $350–$500 hourly rate | $7,000–$15,000 | $84,000–$180,000 |
| Fractional AI lead · about 2 days/week | $8,000–$25,000/mo · rate card published June 14, 2026 from a self-reported survey of 68 consultants | $8,000–$25,000 | $96,000–$300,000 |
| Agency retainer · practitioner framing | About $5,000/mo · a solo advisory practitioner’s own comparison figure, published on a page selling a course on this topic | $5,000 | $60,000 |
| Adjacent titles that are a different job | |||
| “AI Operations Manager” · MLOps framing | $115,000–$185,000/yr · job-description reference site, dated May 16, 2026, describing production-ML oversight | $9,583–$15,417 | $115,000–$185,000 |
| “AI Operator” gig listings | $19–$25/hour · general job-board hourly rate for labelling and monitoring work, at a notional 2,080-hour year | $3,293–$4,333 | $39,520–$52,000 |
Three things fall out of that table once the numbers are on the same axis. First, the employ rows span $75,000 to $247,000 — a factor of about 3.3, or $172,000 of spread for what the titles present as one job. That is not a market failing to price a role. It is at least three roles: a monitoring job, a cross-functional ownership job, and a senior enablement job inside a company where AI is the product. The spread tracks the reporting line more reliably than it tracks the title.
Second, fractional is not the budget option. The cheapest individual fractional retainer in the set annualises to $84,000, which is $9,000 more than the floor of the cheapest full-time band — before you load a single employer cost onto the salary side. Load them and the two routes cross, which is the honest version: at the bottom of both ranges the costs are broadly comparable, and you are buying speed and senior judgement rather than savings. Only at the top does the picture invert cleanly. A $25,000-a-month retainer annualises to $300,000, more than the highest posted base salary anywhere in the table.
Third, the rate cards are at least internally consistent, which is faint praise but worth checking. Four hours a week is about 17.3 hours a month, so a $7,000 floor implies roughly $404 an hour; eight hours a week at the $15,000 ceiling implies about $433. Both sit inside the $350 to $500 band the same source quotes. The arithmetic holds. Whether anyone in your market actually charges it is a separate question, and one you answer with two phone calls rather than a blog post.
05 — The DecisionHire, grow, or go fractional — read your own signals.
Every source we found writes this as an enterprise problem. At twenty to two hundred people you cannot fund a $200,000 enablement lead, so the real choice set is narrower and more interesting: hire someone junior into the band you can afford and grow them, formalise the person already doing the work, or rent senior judgement by the month. The table below maps the signals we actually see in companies of that size onto those three routes.
Read down your own column of signals first, then across. Most companies match on two or three rows, and when they do the rows usually agree.
| Signal in your company | Hire full-time | Grow someone internal | Fractional or retainer |
|---|---|---|---|
| Team signals | |||
| Someone is already doing the work unofficially | Don’t. Hiring over the person already doing it is how you lose both of them. Post externally only after they have declined the role. | Strongest case on the board. Formalise the title, fund the time, and re-band the salary before an external offer does it for you. | Useful as a coach for that person, not as a replacement. Two days a month of senior review, not two days a week of delivery. |
| Three or more agents running with no single owner | Justified once the agents touch revenue or customers. This is the clearest hire signal in the list. | Works if the candidate already owns one of the three. Add 20% of their week and take something off their plate in writing. | Good for the first inventory and ownership pass; poor for the ongoing incident load, which does not respect a retainer calendar. |
| One department accounts for most of the AI use | Premature. A company-wide role that reports nowhere near that department will be politely ignored by it. | Right answer. Promote inside that department, widen the remit only when a second department asks for help. | Only when that department is the one you cannot staff — regulated finance or clinical operations, typically. |
| Budget and risk signals | |||
| No budget above roughly $2,000 a month | Out of reach. The lowest reported band still implies about $6,250 a month in base salary alone. | The only funded route. Budget the training and the tooling, and protect the calendar time — the time is the actual cost. | Also out of reach. The cheapest published retainer in our source set is $5,000 a month. |
| Regulated or sensitive data in scope | Preferred. Accountability that survives an audit is hard to buy by the hour. | Works when the person already carries the compliance relationship. Often they do, and that matters more than AI fluency. | Use for design review and policy drafting. Keep the named accountable owner inside the company. |
| A pilot is approved and the production budget is not | Wrong order. Decide who owns the thing after the pilot before the pilot starts, or your new hire inherits a corpse. | Name the owner now, even at 10% of a week. Naming costs nothing and is the step most often skipped. | Fine for the pilot itself. Write the handover into the statement of work, not into the closing conversation. |
The last row is the one that costs companies the most money. The consulting literature we reviewed puts roughly one in three AI pilots dying in the handoff from pilot to production, for a boringly structural reason: the production budget was never set aside before the pilot was approved. A pilot with no named owner on the other side is not a pilot. It is a demo with a longer runway, and the decision about who runs it afterwards has already been made by default.
One more framing worth keeping. A solo advisory practitioner writing for small-business owners frames the choice as a staff member from around $70,000 a year against an agency retainer of about $5,000 a month — which annualises to $60,000, so the retainer lands about $10,000 below the salary before employer costs. He is selling a course on this exact subject, so treat it as an illustrative contrast rather than market guidance. The useful part is the shape: at the entry level these routes cost roughly the same, and the decision is therefore about what you need, not what you can afford.
06 — OnboardingThe first ninety days: one workflow, all the way through.
A published practitioner structure for the first ninety days in this role is worth reacting to, both because it exists and because it is unusually disciplined about scope. Again, this is attributed analysis from a secondary source rather than a validated methodology — but the constraint at its centre is the right one. One workflow, taken all the way to a measured operational baseline, before anything else starts.
Diagnose
Audit the informal and shadow AI use that already exists, interview three to five people across different teams, and pick exactly one workflow using explicit criteria: volume, repetitiveness, time cost, stakes, and whether the data is actually reachable. The audit is the part people skip, and it is where the surprises live.
Build and validate
Write the system prompt, build a ten-case test set, and pilot with two or three users only. Ten cases sounds small until you try to write them honestly; three users sounds small until one of them finds the failure mode that would have embarrassed you in front of forty.
Stabilise
Establish a pass-or-fail operational baseline, connect one live data source, and deliver a one-page report to the person who funded this with a real number in it. Only then does workflow number two begin. The one-page report is what buys the second quarter of budget.
The same source is equally clear about what to defer, and this is the part we would underline for any company under two hundred people. Skip building a custom protocol server — native connectors and general automation tools cover the large majority of first-year needs. Skip the custom monitoring dashboard; a spreadsheet and a weekly thirty-minute review is enough at this stage. And skip company-wide training until three agents have been proven reliable, because training people on an agent that breaks in front of them destroys trust and, per that analysis, sets adoption back by months.
Add one item to the list that the source does not include: the register. From day one, the person in this role should own a single list of every agent running, who owns it, what data it touches and how to stop it. The internal agent registry template is the artifact we would hand them on their first morning. It takes an afternoon to populate and it is the thing that makes the incident duty survivable.
07 — Failure ModesThe dominant failure is not a bad hire — it is an unfunded side-quest.
Two named failure modes appear in the published ninety-day material, and a third one appears in nearly every small company we have watched attempt this. The first two are about how the person works. The third is about how the company set the job up, and it is by some distance the most common and the most expensive.
Chaos
Trying to automate five workflows across three teams simultaneously. Nothing reaches a reliability bar, every team gets a half-working thing, and users quietly stop using the output. Recovery is slow because trust decays faster than it rebuilds.
Paralysis
Spending the entire first quarter on architecture diagrams and stakeholder alignment with nothing in production. The sponsor’s enthusiasm has a half-life, and it is shorter than ninety days. The second budget conversation is the one that never happens.
The unfunded side-quest
The work is handed to one already-loaded engineer with no budget line, no written scope, no time removed from their existing job and no ninety-day plan. It is not a hiring failure, because nobody was hired. It is a scoping failure, and it is invisible until the person leaves.
The tell for the third one is always the same sentence in a status meeting: “Yes, so-and-so has been looking at that.” Nobody has written down what so-and-so is accountable for, what they are allowed to spend, or what they were told to stop doing to make room. The work then competes with their actual objectives, loses, and resurfaces three months later as a production incident nobody has time to fix.
The fix is embarrassingly cheap relative to the cost of the failure. Write down the scope in a paragraph. Put a number against it, even a small one — a tooling budget of a few hundred dollars a month is a different signal from no budget at all. Remove something from their existing workload explicitly and in writing, so the trade is visible to their manager. Set the ninety-day target as one workflow at a measured baseline, not as a portfolio. And name the person in a document that other people read, because ownership that only exists in a hallway conversation is not ownership.
08 — The PostingWriting the ad: lead with the reporting line, not the buzzword.
The framing you lead with determines the applicant pool far more than the salary does. Three framings work at this size, and they attract genuinely different people. Pick the one that matches where the role reports and what your first ninety days actually need, then name the alternate titles in the body of the ad so search picks them up.
AI Ops Manager
Lead with reliability, process and cost per workflow. Attracts operations people who will instinctively build the register, the review cadence and the escalation path. Weakest on hands-on prompt and integration work, so pair with an engineer for the first two builds.
AI Enablement Lead
Lead with adoption, workflow assessment, governance alignment and training. Both dated live postings we priced used this title, so it is the most searchable of the three. Best when the hard part is people rather than plumbing — which, at most companies, it is.
AI Agent Manager
Lead with systems, integrations, evals and model-upgrade drift. Attracts the most technical pool and prices highest of the three. Justified when agents already touch production systems and credentials, and over-specified when they do not.
Four things to put in the ad regardless of framing. State the reporting line explicitly, because candidates for this role are reading it as a signal about whether they will have authority. State the budget the role controls, even if it is small — it separates a funded role from a hopeful one, and the good candidates have already been burned by the difference. Describe the first ninety days as one workflow to a measured baseline, which screens out anyone who wants to spend a quarter on strategy. And ask for domain experience in the function being automated ahead of AI credentials, because the evidence in the postings points that way and the certification market has moved faster than the competence behind it.
One thing to leave out: any implication that the person will build the platform. They will not. They will configure, connect, test, supervise and teach, and the job is closer to a well-run systems administration function than to engineering. Once you are running more than one or two agents, the operations-team playbook for process automation covers what the function looks like when it grows past a single person. And if the honest answer to “who owns this?” is still nobody after reading this, that is the actual first task — our AI and digital transformation engagements typically open with an ownership-and-inventory pass for exactly that reason.
09 — ConclusionName the job before the incident names it for you.
The title is unsettled. The work is not, and somebody is already doing it.
Every company running agents has this job. The only variable is whether it has a name, a budget and a plan, or whether it is being absorbed into someone’s evenings. The published pay data spans about 3.3 times from floor to ceiling because the market is pricing at least three different jobs under one label — and the useful consequence for a company of twenty to two hundred people is that you get to choose which of those three you are actually buying.
Our read of the direction of travel: the titles consolidate, the certification market gets noisier and less informative, and the durable signal stays what it is today — somebody who has run agents somewhere, in any function, under any title. Domain knowledge plus demonstrated operating experience will keep beating credentials for at least another hiring cycle, because the hard parts of this job are judgement calls about the business rather than facts about models. The one thing we would bet against is the standalone AI ops headcount surviving at small companies. It gets absorbed back into operations or IT once the workflows are stable, the way webmaster did.
The practical next step costs an afternoon. Write down every agent running in your company and who touched it last. If more than one name appears in that second column and none of them is accountable in writing, you have found the job — and you have almost certainly also found the person who should be doing it. Formalise them before the market does.