Insights
The first 90 days of AI in an understaffed city IT department
A 90-day AI plan for city IT: what to ship first, what it actually costs, what to tell the council, and what the CPRA means for your prompts.
NASCIO’s State CIO Top Ten for 2026 put artificial intelligence at number one for the first time in the list’s 20-year history, ending cybersecurity’s 12-year run — in the same season the National League of Cities reported general fund revenues declining 1.9% in FY2025. If you run IT for a city, that pairing is now your job description: adopt the new thing with less money and fewer people than you had last year.
This is the plan we would run in your seat. It names a first project, publishes real running costs, answers the California records-law question nobody else touches, and tells you when to skip AI entirely.
Why is AI suddenly your problem when your budget just shrank?
The adoption numbers moved fast. A May 2026 Tyler Technologies survey reported by ICMA found 44% of local governments experimenting with generative AI, 33% with a defined AI policy, and 30% actively writing one. Two years earlier, ICMA’s own survey had 48% of local officials calling AI a low priority.
The capacity numbers did not move with them. In that same Tyler/ICMA survey, 63% of municipal respondents named limited internal expertise or staffing as the biggest barrier to adopting or scaling AI. MissionSquare’s 2025 workforce survey found 40% of governments that recruited for IT still rate those positions hard to fill, and 46% expect the biggest retirement wave is still coming.
Money is tighter too. NLC’s City Fiscal Conditions 2025 shows general fund spending growth collapsing from 7.5% in FY2024 to 0.7% in FY2025, with only 45% of city finance officers optimistic about meeting FY2026 needs — down from 64% a year earlier.
And your staff did not wait for a policy. MissionSquare’s survey of 2,000 state and local employees found 46% already use AI at work, while 62% have received no employer training on it.
The question in front of you is not whether your organization uses AI. It is whether anyone is directing it.
What should a city IT department ship first?
One internal-facing workflow with a measurable baseline. For most cities that means helpdesk ticket triage — categorize, prioritize, and route incoming tickets before a human touches them — or agenda and minutes summarization for the clerk’s office. Not a chatbot, not a strategy document, not a committee.
The reason is risk asymmetry. An internal tool fails in front of staff who can flag the error and move on. A public-facing tool fails on the record, in front of residents and reporters, and the failure becomes the story.
The internal-first pattern has real precedent. San Francisco rolled Microsoft 365 Copilot Chat out to roughly 30,000 employees in July 2025 only after a six-month pilot in which 2,000-plus employees reported saving up to five hours per week on email drafting, planning, and summarization. At the other end of the size scale, the clerk in Long View, North Carolina went from replaying meeting recordings in 15-second increments to AI-drafted minutes — about five hours saved per meeting, per the vendor’s own figure.
Scope it like an operator: one workflow, one owner, one before/after metric. Minutes per ticket or hours per minutes packet, recorded before the pilot starts.
Most first projects need an assistant, not an agent — the distinction matters when vendors quote you, and the same taxonomy applies to city workflows as to SMBs. Our use-case catalog lists the triage and summarization patterns we see repeat across sectors.
Why not start with a resident-facing chatbot?
Because New York City already ran that experiment for you. The MyCity chatbot, launched in 2023, told businesses to break the law — that landlords could refuse Section 8 vouchers, that employers could take workers’ tips, that restaurants could refuse cash. The city kept it live as a “pilot” with disclaimers until Mayor Mamdani called it “functionally unusable,” put its cost at roughly half a million dollars, and had it taken down by February 4, 2026.
The public-facing deployments that do work reveal the hidden cost. Denver’s “Sunny” chatbot was handling 12% of all 311 inquiries in 72 languages as of mid-2024 — and has been documented linking residents to stale 2019–2021 content, because a chatbot is only as current as the knowledge base someone maintains. Sacramento’s 311 AI transferred 52,298 calls in FY2024/25, a real result — behind a 311 operation with dedicated staff.
Resident-facing AI is a year-two project. It assumes a maintained knowledge base, a comms plan, and someone on the hook when it says something wrong.
The tradeoff is real: you give up the visible win the mayor wants in the press release. The council memo section below covers how to sell the boring project instead.
What does municipal AI actually cost to run?
Nobody ranking for these queries publishes a single number, so here are sourced figures at three tiers.
| Tier | Example | Sourced cost | What the invoice omits |
|---|---|---|---|
| Per-seat copilots | M365 Copilot-class tools | Enterprise list pricing ~$30/seat/month; Microsoft added a cheaper under-300-seat tier in late 2025 | Rollout training, prompt guidance, usage review |
| Single-purpose tools | San Jose Wordly translation; Denver Citibot | ~$82K/yr vs ~$400K for human interpreters; $184K year one, ~$100K/yr after | Knowledge-base maintenance hours, content freshness audits |
| Enterprise plan review | Austin Archistar | $3.5M over 3 years | Human final review on every output |
The per-seat math is simple enough to run yourself. At roughly $360 a seat per year, a tool pays for itself if it saves a $60/hour loaded employee six hours annually — half an hour a month.
San Francisco’s pilot reported gains of up to five hours per week; even if your staff captures a tenth of that, the seat clears its cost many times over. The honest caveat: “up to” is doing work in that sentence, which is why you measure your own baseline instead of trusting anyone’s pilot.
The Austin number carries a second lesson. Its pilot found the AI accurate about three-fourths of the time, struggling specifically with Austin’s tree and flood rules — and Austin still bought it, because a 75%-accurate first pass with human final review beats a zero-percent first pass. Even mature vendors need a human at the end of the line.
Budget the line items vendors leave out: setup hours, monthly prompt and knowledge-base maintenance, and evaluation time to check outputs against ground truth. On the automation layer underneath these tools, the build-vs-buy cost logic is the same for a city as for a business, and a scoped automation engagement should quote all three omitted line items up front or you are not looking at the real price.
Are AI prompts and outputs public records in California?
Probably yes, and you should operate as if the answer is simply yes. The CPRA defines a public record as any writing relating to the public’s business that is prepared, used, or retained by an agency. Per an April 2026 analysis by Liebert Cassidy Whitmore, a California public-agency law firm, that definition plausibly captures prompts (prepared by employees), outputs (used in agency business), and vendor-hosted chat logs (retained, even on a third party’s servers).
This is not theoretical. The same analysis notes journalists have already obtained officials’ full ChatGPT conversation logs through records requests — a Cascade PBS investigation in Washington surfaced AI-drafted constituent emails and mayoral letters — and at least one California city has received a CPRA request specifically seeking “generative AI records.” San Jose said it plainly in its July 2023 employee memo: presume anything you submit could end up on the front page of a newspaper.
The practical guidance is short. Write prompts the way you write emails to the council.
Decide how AI chat logs map to your retention schedule before deployment, not after the first request lands. And ask every vendor two questions in procurement: where do chat logs physically live, and who responds when a CPRA request arrives?
None of this is legal advice — it is an operator’s read of a law firm’s analysis. Put thirty minutes with your city attorney on the calendar before the pilot starts.
What data can never go into a commercial LLM?
Start with police data, because the answer there is absolute. Anything touching Criminal Justice Information falls under the FBI CJIS Security Policy: contractors must sign the CJIS Security Addendum, compliance is administered state by state, and personnel are subject to the screening requirements of the FBI’s CJIS Security Policy, including fingerprint-based background checks.
Per compliance-industry analysis from FirmAdapt — analysis, not a primary source — California’s CJIS Systems Agency treats a model trained on CJI as CJI itself, and major commercial LLM APIs have no published CJIS addendums. So: no PD data in general-purpose chatbots, full stop.
HIPAA reaches further into city hall than most IT directors expect. Cities become covered or hybrid entities through EMS operations, health clinics, and self-funded employee health plans, per UNC’s Coates’ Canons. The working model is San Francisco’s July 2025 generative AI guidelines, which gate protected health information to tools operating under a Business Associate Agreement — no BAA, no PHI, no exceptions.
Round out the checklist with utility-billing PII and personnel records. Neither belongs in a free consumer chatbot under any policy you would want read aloud at a council meeting.
The workable alternative is retrieval grounded in your own tenant. A RAG system scoped over your documents keeps the data inside your existing Microsoft or Google environment instead of shipping it to a third party. But apply the honest threshold we give businesses: under roughly 500 documents, RAG is usually not worth building, and the same math applies to a small clerk’s office as to a small company.
What do I tell the council — and the union?
Write one page, and pre-answer the four objections you already know are coming.
Jobs. Austin publicly committed that no positions would be eliminated when it adopted Archistar. AFSCME’s bargaining guidance instructs locals to negotiate before deployment decisions and to demand information from both the employer and the vendor, and San Jose unions are pushing contract language that the city may not use AI to replace workers. The move is obvious: bring the union in during week one, not after procurement.
Privacy. Point to the data-boundary checklist above and the CPRA posture: nothing sensitive enters a commercial tool, and every prompt is written as a public record.
Accuracy. Human final review on everything, with Austin’s 75%-accurate pilot cited as the industry norm rather than a scandal. You are buying a fast first draft, not a decision-maker.
Cost. The per-seat break-even math from the cost section, with your own baseline metric attached.
On procurement, small dollars have paths around a full RFP. Cooperative vehicles like Sourcewell’s competitively solicited contracts and California’s CMAS schedules, open to local agencies, let you piggyback on a solicitation another agency already ran — confirm what your own purchasing thresholds allow with counsel. Watch the state layer too: California’s March 2024 GenAI procurement guidelines require GenAI disclosure language and written solicitations at the state level, and Executive Order N-5-26 (March 2026) is adding vendor safety certifications and data-minimization contract provisions that will trickle down to city purchasing.
Timing is the constraint nobody mentions. California cities run July 1–June 30 fiscal years with budgets adopted in June — Campbell’s cycle starts in December — and some, like Alameda, budget on two-year cycles. An unbudgeted purchase waits for the next cycle; a pilot small enough to fit existing appropriation is the practical move.
What does a realistic 90-day plan look like?
We call it the AIX 90-day municipal pilot plan, and its spine is measurement — anecdotes are how the thin content ranking above this piece got written.
Days 1–30: policy, inventory, baseline. Adopt a use policy by adapting the GovAI Coalition templates — the coalition San Jose founded now spans more than 3,000 members across roughly 900 agencies and is spinning out as an independent nonprofit, making it the de facto standards body. Survey staff to inventory the shadow AI use that is already happening. Pick the one workflow and record its baseline: minutes per ticket, hours per minutes packet.
Days 31–60: run the pilot. Three to five staff, the one workflow, a shared log of every failure and every save. Hold a 15-minute review each week — the failures teach you more than the wins, and they become your eval set.
Days 61–90: measure and decide. Compare against the baseline you recorded, write the one-page council memo from the section above, and make a call: scale, kill, or extend. All three are legitimate outcomes; the only failure mode is drifting into month four without a decision.
If you want the shortcut to step one, our AI Readiness Assessment covers the same four dimensions — data readiness, workflow maturity, team capacity, use-case clarity — in ten questions.
When is AI not worth it for a city IT team?
More often than any vendor will tell you. Skip the pilot for now if any of these describe you:
- Your records are paper or scattered PDFs with no digitization budget. Fix the data first; that is the Foundation tier of our assessment, and no model rescues inputs it cannot read.
- Ticket volume is low. If triage saves minutes a week rather than hours, the maintenance overhead exceeds the win.
- Nobody owns knowledge-base upkeep. You will re-create Denver’s stale-content problem, except without Denver’s 311 staff behind it.
- The workflow touches CJI or PHI and you have no compliant tenant path. The checklist above is a wall, not a speed bump.
- The budget covers the license and nothing else. No line for maintenance hours means the tool degrades quietly until someone notices in public.
There is also a case where the fix is policy, not procurement. If the real problem is untrained staff pasting resident data into free chatbots — remember, 62% of government employees have had no AI training — then a use policy and one training hour beat any purchase you could make this year.
Invert the list and the conclusion writes itself: if none of these apply, the 90-day plan is low-risk enough to start this quarter inside existing appropriation.
Where does outside help fit for a one-person IT shop?
The honest case for outside delivery is the 63% of municipal leaders who name staffing as their biggest barrier. The constraint is hours, not ideas — most IT directors we talk to could sketch the right pilot on a whiteboard; they cannot staff days 1 through 90.
A scoped engagement should look like this: the outside shop builds the pilot and the eval harness, your team runs the weekly reviews, and by day 90 your team owns the whole thing. That is how we scope agent delivery, and being a San Mateo shop, we operate under the same CPRA, CJIS, and June-budget-adoption constraints as our neighbors — which is why this playbook is deliberately California-specific.
One thing worth saying out loud: most of the cities in the surveys above are doing this without consultants. This playbook is complete enough to run alone, and that is by design.
Keep reading
Related
Want this shipped in your org?
Take the AI Readiness Assessment for a personalized scorecard, or book a 30-minute discovery call. No slides — bring the workflow you've been arguing about and we'll tell you whether it's the right first one.