The proposal here is to make evaluation a loop rather than an event. Multi-stage and longitudinal evaluations accompany a chatbot across its lifecycle, each round's findings graduating into consistent validators, and the risks the deployment documents gathering into risk cards, living reports that carry findings into every round that follows.
Are evaluations of social-sector AI chatbots often blind to deployment stage? Do their findings stay in reports rather than become guardrails? I grappled with both by screening thirty-four sources under a written protocol. The reading exposed two structural holes in the framework field. Nothing in it tracks a deployment across its lifecycle, a norm the literature itself states (Weidinger et al., 2023). And I am yet to come across specific frameworks in the field that connects evaluation findings to the artifacts that would enforce them. Hence the loop below. Tattle's four-step manual loop is the engine. Lifecycle stage gates it. Per-language human gold sets constrain it, because judge reliability collapses exactly where these deployments live (LLM Safety Alignment in Low-Resource Languages; Fairness or Fluency?). It ends in artifacts that keep running: guardrail validators (Guardrails × MLflow) and Social-Sector RiskCards (RiskCards).
# Where this began
This piece began as part of my attempt to understand Tattle, an Indian civic-technology organization that builds tools and datasets to understand and respond to inaccurate content including the chat platforms where much of India's public conversation happens. As part of that attempt, I was asked to share my thoughts on the risks of AI chatbots in social sectors: against what Tattle's approach has been and how they are thinking about it, which frameworks I would rely on, and how I would evaluate these systems. What follows is that answer taking shape. I read what Tattle reads, ran a review of the wider literature, and and arrived at a framework I believe in.
Chatbots are moving into public-service delivery in health, education, legal aid, and welfare, run by nonprofits, government programs, and development organizations. The documented misuse arrived with the deployments, hardly before: bots fielding sex-determination questions; risky queries reframed as study questions to slip past filters; an admissions bot whose misuse surface required a stress test before launch. Not hypotheticals. Launch records. Following Weidinger et al.'s sociotechnical harm taxonomy, which sorts harms by the model's capability, the human interaction around it, and the systemic impact of both, I read these instances as clustering into three layers: user harms, organizational harms, societal harms. The risks section later in this piece turns these layers into a risk × evidence table of documented social-sector cases. That table carries the argument's weight. Let me only set up the layers here.
Are we treating AI safety like web accessibility? The ARIA labels that make interfaces usable for disabled users, de-prioritized, added if time and money remain. Cybersecurity, by contrast, is a precondition for shipping. When I read the field map by Baarish, I found the literature had already drawn the same curve: biases become "part of the model's knowledge system," from-scratch multi-stakeholder redevelopment is something corporates will not fund, so the guardrails arrive in hindsight, under public pressure (What is AI Safety?). Baarish joined Tattle in November 2025 as Research Development and Grants Specialist and wrote both "What is AI Safety?" and the manual-evaluation guide this article leans on repeatedly. I would rather build the guardrails before the pressure. That is the whole argument for acting now.
The structure behind the urgency, as I came to see it: a consumer product with a bad answer loses users, and the market corrects. A public-service bot has no such correction. That is also why I do not believe evidence about chatbot risk gathered on English-language consumer products transfers unexamined: neither the misuse patterns, nor the correction mechanism, nor the language carries over. The documented misuse arrived with the deployments, and it arrived aimed at systems that hold people's health, education, and entitlements, not their attention. Underneath, the interaction is information-asymmetric. People consult the bot for exactly the knowledge they lack, and are neither supposed nor expected to verify its answers. And the obligations differ, which the retention framing hides. A consumer product owes its user nothing but a working product they are free to leave; keeping them is the whole discipline. A public-service bot owes accuracy whether or not anyone stays: legal obligations in service mandates and liability, ethical obligations of care to people who are assigned or enrolled or simply dependent, and the plain accuracy a wrong answer violates. A market corrects by exit. A public service must correct by duty, and nothing in a consumer-product evaluation ever tests the duty.
One observation from the search run itself doubles as gap evidence. Between a full-time job and the standing obligations of family and friends, I did not have the hours to survey the grey literature properly myself. What little time I had went into deploying research agents carefully: a written plan, a protocol, then the searches. What came back was almost as informative as what did not. Searches naming NGOs and nonprofits returned essentially nothing on arXiv; the social-sector chatbot literature barely exists there, one irrelevant hit across the development-sector searches. It lives in practitioner conference venues like Information and Communication Technologies for Development and community-centered human-computer interaction, and in organizational blogs. Even "AI safety Global South" returns only two preprints, both weak fits for chatbot risk; the actual cluster lives in the grey literature. arXiv was the one preprint venue the agents covered this time; next time, if time permits, I want to loiter on SSRN and a few of the other platforms where practice writes first and papers later. The actual cluster lives in the grey literature. The discourse this article joins is grey-lit native.
# How I read the papers room
I am an engineer first, before I am an evidence synthesist. When I turned from a codebase to a body of literature, the engineering way of working took over the evidence-synthesis process and hybridized it. Spec first: the research questions, the inclusion and exclusion criteria, the stopping rules, the annotation scheme, all written down before the first search ran. Then a small, time-boxed scoping review, run as a loop I trust until it produces something defensible.
The read set, each item annotated through dictated voice notes: Tattle's AI-safety services page; the manual-evaluations guide, with a first-hand look at the slur list from Uli, Tattle's dataset and tool for online gender-based violence; What is AI Safety?; Towards Functional Safety; Tattle's co-founder and research lead Tarunima Prabhakar's Tech4Dev guardrails reflections, where the writing anchors this corpus's deployment-side thinking; the Uli dataset paper; the Kaapi Guardrails build story; the Guardrails×MLflow tooling post; and of course Weidinger et al. 2023 as the one deep read. I annotated live while reading, dictating into the addictive Wispr Flow. The annotations simply flowed, and reading my thoughts out loud kept me honest in a way silent note-taking never managed. That silence had its own treachery: it censors the flow at the keyboard, or hands a thought a structure it has not yet earned. Dictation permitted neither, which is a way of saying the blabbering was the thinking. Either way, the scale surprised me: just under 19,500 words of dictated annotation across eight voice-note files, the last of them a closing analytic memo. And I only noticed after the fact that the loop I used to study the literature is the same loop the literature itself recommends: sample the field, annotate what I sampled, expand on the edges, analyze, then argue from the annotated evidence for the one gap worth working on.
The mechanics, in plain words, and the agents ran most of them. The opening section owns the why: the hours to survey the grey literature properly myself were not there, and carefully deployed agents were the middle path between doing it badly alone and not doing it at all. The split held: the brief and the screening calls were mine; the seed strings, the searches, and the citation chasing were the agents'. Briefed with my annotated readings, they planned fifteen search strings aimed at the four kinds of place a researcher looks: academic databases, standards and policy bodies, practitioner and grey literature, and evaluation-tooling documentation. They traced citations forward too: who cited the Uli dataset paper and who cited the AILuminate safety benchmark, because a field's citing literature is where it says what it did with its own tools. One planned seed, the manual-evaluation lineage, I skipped deliberately, because the assigned corpus already owned that ground. The run produced 44 screened items, and every one carries a one-line decision: 34 included, 9 excluded with reasons, 1 already in my assigned corpus, zero silent discards. Four of the included became must-reads. The corpus froze on the protocol's stopping rule, three consecutive well-formed searches adding nothing new. The rule fired when the AILuminate citing literature had turned into benchmark-usage, jailbreak, and domain-guardrail work with nothing new on social-sector chatbot risk: 33 titles screened to a bulk no in a single pass. One scope limitation stays on the record: the review is English-first. Hindi and other Indian-language sources were chased only when they surfaced through footnotes and citations, so the local-language grey literature is under-covered by construction.
The searching was the agents'; the deep reading was mine, and no agent could have had it anyway: the urge to read, to own, to understand does not delegate, and mine would not quiet. I read compulsively, racing the clock for every hour. A few articles, but read the slow way: aloud, annotated, argued with. What does the paper say? What does the paper name: risks, methods, deployment stage, context, evidence type, outputs? Can I trust it, and what does it assume? And, unedited, what did it trigger in me. Under time pressure, the research agents helped press my raw dictated annotations into these four layers. An analytic memo closed each reading session: what changed, what surprised, what now looks wrong. Two memos exist, and they feed the frame of this article. I wish I had had more time for more deep readings. The Weidinger sociotechnical paper, the one must-read I read fully, grabbed me and left me very impressed. More than that, it changed my attitude toward this whole exploration: what had been a vague, broad chase after a moving target took on a clearer, well-structured shape, and that structure added confidence and meaning in the exploration, and in its possibilities and potential.
Out of curiosity I asked my research agent whether this improvised approach would validate against any named methodology. It returned three: thematic synthesis (Thomas and Harden), framework synthesis (Brunton et al.), and the scoping-review sensibilities of Arksey and O'Malley. I still owe those abstracts a proper verification read. Noted for later. I need to finish writing this piece first. But for now, I take the names as retrospective comfort that I was not improvising blindly, not as recipes I followed.
# Every bad response is bad in its own way
Tarunima's reflections on building guardrails for the Tech4Dev cohort observe that everyone agrees on what a good response looks like, but each bad response is bad in its own way, needing its own diagnosis. No single rubric grades every failure, so I went looking for where each failure has already been documented. The documented instances cluster in three layers, and the cleanest way I found to hold them was Tattle's own documented engagements rather than the taxonomy's abstractions. The Tech4Dev cohort collaboration included a nutrition bot and an education bot, and both document the user layer: prescriptive out-of-scope answers, a caste-category disclosure during routine profile queries, and language-switching failures that exclude the very users the bot serves (Manual Evaluations of AI in the Social Sector). The Kaapi Guardrails build showed me how the organizational layer absorbs the user layer's findings. Evaluation findings turn into validators, and the validators themselves need maintenance. The admissions-bot stress test for IITM (Indian Institute of Technology Madras) is where the organizational layer appears in its own right: trust erosion, liability, a misuse surface red-teamed before launch (Tattle's AI-safety services page). User harms have their engagement stories; organizational harms have theirs. The societal layer has none yet. And that no engagement story has surfaced this layer is itself a finding. The table below is this article's evidence base; each row anchors a risk class to its documented instance.
Then the coverage read, which drives the rest of the article. The table above holds the risks with documented instances; the fuller map sets sixteen risk classes against the four lifecycle stages, pre-launch, pilot, production, and longitudinal, meaning tracked across time after deployment, and asks one question of every intersection: does the corpus contain evidence of anyone testing for this risk, at this stage of the deployment's life? Values came from the reading notes: each source carries a structured extraction of risks named, deployment stage, and evidence type. A cell fills on any deployment-level evidence, counts partial when it is indirect, model-level only, or adjacent domain, and stays empty otherwise. Sixteen risks, four stages, 64 answers: 17 filled, 16 partial, 31 empty.
Coverage clusters where evaluators are actually in the room, pre-launch and pilot, and thins once real users arrive: production is the second-weakest column, and the longitudinal column is empty in all sixteen rows. Nobody in this field has followed a deployment past launch. The blankness is the finding, and it is the same finding as the title: evaluation leaves with the evaluators. An engineer reads an empty matrix the way they read an untested code path: nobody has hit the bug yet, which is not evidence the code works.
One mechanism correction of my own, because it changes what the fix looks like. The caste case is not the model repeating something memorized from training data. It is in-context profile surfacing. The bot echoes a user's own caste category back at them, information the conversation itself supplied, now written into logs. I file the case under ProPILE's probing taxonomy for personal-information leakage, with the caveat that the taxonomy never anticipated caste markers. A later study complicates the leakage story, arguing apparent leakage may be cue-driven rather than memorized, but that contradiction never touches the caste observation, which is about context and logs, not model memory. The larger point stands: standard personal-information taxonomies are silent on caste markers.
# The framework field and its two structural holes
I placed every evaluation approach I had read, twelve of them, from NGO playbooks to NIST risk frameworks, side by side and asked each the same nine questions: which stage of a deployment does it cover, who defines harm, what does it leave behind, and so on. Two holes fell out of that grid, and they drive everything I propose next.
The first hole: nothing I read tracks a deployment across its life. The stage question drew the weakest answers of the nine. The only approach that even gestures at it, the Agency Fund's four-level stack, turns out on reading to rank different things to evaluate, the model, then the product, then the user, then the impact, rather than different moments to evaluate them. And it expects all four to be checked continuously. Weidinger et al. prescribe evaluation "at multiple time points throughout the AI system life cycle, including by monitoring effects post deployment", and I found no framework that turns that sentence into a practice.
The second hole: Evaluations end exactly where guardrails begin. Look at what people ship and the field falls in two. Half produce scores and reports. Half produce running artifacts, datasets, validators, test suites. Between the two, nothing. There is no framework for turning a finding in a report into a check that keeps gating the system after the evaluators leave.
The bridge between what the norm says and what I build in the next section runs through two old ideas and Weidinger's own conditions. Collingridge's dilemma says that early in a technology's life we can still steer it but cannot see it clearly, and that late we can see clearly but can no longer steer. That is why the pre-launch gate and the longitudinal gate are different instruments, not the same gate repeated. Early evaluation shapes direction; late evaluation reveals accurately. Goodhart's development-versus-assurance separation says the loop that improves the product must stay distinct from the loop that judges it, which is why a frequent regression batch and a deep annual re-evaluation are not the same activity. And Weidinger supplies the conditions' teeth: evaluation "derives its relevance from the processes and decisions into which it is meaningfully embedded"; and facing the trade-off between longitudinal, localized accuracy and the generality of automated tests, the paper concludes that "establishing a mixed-methods practice is the way forward".
Models are less safe in low-resource languages, and the tools that would measure that, LLM judges grading other models' outputs, are least reliable in exactly those languages. Pairwise judges prefer English answers regardless of correctness and do worst on culturally grounded subjects Fairness or Fluency?. The field's first systematic review of the safety gap LLM Safety Alignment in Low-Resource Languages prescribes participatory, community-defined evaluation as the fix. Tattle already works that way: Uli's survivor-defined annotation, the manual eval loop, the MLCommons expert group. The literature hasn't caught up to its own diagnosis, so per-language human gold sets are non-negotiable.
My own reservations sharpen this rather than weaken it. The three-layer taxonomy, Weidinger's division of evaluation into capability, human interaction, and systemic impacts, may not be that helpful for Tattle. From my read it does not add enough details, and a framework lacking both simplicity and nuance ceases to have utility, momentum, direction. For Tattle's goal we may have to look at other frameworks, or possibly come up with one. The paper's argument that component-based safety must give way to systems thinking feels like a component-based approach itself, because the components are the model, the context, and the interactions between the two. On who defines harm, my position is definitions set closest to the harm, community-level, with institutional layers as checks, legal floor and platform ceiling, never as sources. Tattle's guide already practices it, co-developing in-scope and out-of-scope definitions with the client. And one worry I never resolved, named after Heisenberg's observer effect: the act of measuring a deployed model changes what you know and what you must re-measure. Models are always evolving; you measure a moving target, and it is position versus speed, so measure often. I have no answer, only a design consequence.
One fact outside the matrix, the same story at organization scale, and this one personal. After reading everything, I was still close to clueless, a broad idea with no shape. I badly needed an artifact to make my thinking presentable, and the short micro-paper the assignment asked for felt helpless, because underneath I was craving to write. So I let the writing process itself do the thinking, crammed into little time alongside a full-time job and a weekend with friends, family, and visitors. Still worth it, because the artifact structures the thought, and that is how the framework in the next section got built. In hindsight I could not fully substantiate every choice in it, and with more time I would have derived alternative frameworks with more rigor; so far, so good.
# A stage-gated, context-first evaluation loop
I am an engineer, and one urge shaped where this framework ends. Findings itch. The moment an engineer reads one, the reflex kicks in: turn it into a check, wire it into the pipeline, make it run on every commit. That's the graduation. From evaluation to validation. So every round terminates in running artifacts, checks that keep gating the system after the evaluators leave. But the same urge, undisciplined, would pre-direct the evaluation toward whatever is easy to codify. The design holds it in check twice. The Not-Doing list, a written list of what this framework refuses to build and why, refuses to engineer for unobserved risks. And Goodhart's development-versus-assurance separation keeps the loop that improves the product distinct from the loop that judges it.
The engine is Tattle's four-step loop: sample, annotate, keyword-expand, analyze. The most deployment-tested method in the corpus, proven in three NGO pilots Manual Evaluations of AI in the Social Sector. Four constraints sit around it, each doing one job:
-
Stage gates split the lifecycle into pre-launch, pilot, production, and longitudinal stages. Each with its own test battery, exit criterion, and a match to how much of the system the evaluator can actually see.
-
Language robustness: every automated evaluation is calibrated against a per-language human gold set (a small human-labeled reference dataset), because judge calibration in English doesn't transfer — pairwise judges prefer English answers regardless of correctness and are worst on culturally grounded subjects Fairness or Fluency?; Pradhan et al., 2025.
-
A tier split: deterministic validators catch the known badness, judges take rubric work, humans keep discovery. All complementary.
-
Codification: every round ends in validators plus a Social-Sector RiskCard; re-evaluation measures whether the guardrails worked.
Two builds show the tier split is buildable. Kaapi Guardrails turned evaluation findings into four running validators on the Guardrails AI framework (a slur check, a personal-information detector, a gender-assumption check, a ban list). Guardrails×MLflow runs the same Hub validators as evaluation scorers, turning safety regressions into continuous-integration tests.
The Social-Sector RiskCard deserves its own introduction, because the codification step rides on it. I adapted it from Derczynski et al.'s RiskCards, structured context-conditional risk documentation, reshaped for social-sector deployments. A card carries seven fields: the risk's name; its harm routes, who is harmed and how the harm travels; its placement among the user, organizational, and societal layers above; example prompt-output pairs; harm-definition provenance, meaning who defined this as a harm here, whether survivors, the client, a funder, or the literature; the lifecycle stages where it has been observed; and validator linkage, which running check gates it, or "none yet." A card earns its place alongside the report. The report carries the argument; the card travels with the deployment, converting findings into testable artifacts. It is read at each gate's exit, and the gate closes when every high-risk output class is either gated by a validator or carried on a card whose fields are filled honestly.
One worked card, because the abstraction only earns trust when instantiated. The caste personal-information card, walked field by field:
Presidio is the standard open-source detector for personally identifiable information, and it does not carry caste natively — the Indic entity lists are the extension.
The gates in practice:
- Pre-launch gate. Red-team the misuse surface, the IITM pattern. Localized functional tests via the SGHateCheck recipe: seed tests, LLM translation, native-annotator refinement, ending in a localized Indic suite. Deployment-level personal-information probes with Indic entity lists. Ban-list and slur-list validators installed. Exit. Every high-risk output class is either gated by a validator or carried on a RiskCard.
- Pilot gate. The 4-step loop at the guide's tested ratios, sampling 1.5-2.5% of conversation pairs, spread across time. Opening on day one with the guide's floor: a random sample, a query-type consistency check, a language-consistency check, a personal-information screen. Scope definitions co-developed with the client. Per-language gold sets seeded directly from the annotation round, because the loop's labels ARE the gold set. Exit. Label distributions + close reading delivered as a RiskCard draft, not a report alone.
- Ongoing, production + longitudinal gate. Validators as regression gates, the MLflow scorer form. Judge drift checks against the gold sets at every model update, meaning re-scoring the same reference set to catch a judge silently changing its standards. Small frequent batches + deep annual/bi-annual re-evals, the guide's own build-over-time direction, which no corpus item yet performs. Weidinger states the norm, evaluation at multiple time points and monitoring after deployment, and the practice is absent everywhere I looked.
Three margins to close, each with its resolution.
The first is scale of the artifacts. Each round leaves something behind: labels become gold sets, findings become validators, risks become cards, and round two is cheaper because round one's findings already run as gates. Hence the 4-step loop as engine: a bounded, repeatable process, every step feeding an artifact. The labels stay local.
The second is context. Conversation-length limits push evaluation into context engineering, meaning window length, system-prompt placement, evolving context. Context is where the answers live. Not a practice I was new to, then, just one I had underestimated. The corpus did not introduce context engineering; but it convinced me it carries more than I assumed.
The third is the constructive turn, the one that closes the second structural hole, where evaluations end exactly where guardrails begin. Tattle's Indic harm taxonomies already exist: the annotation labels Uli uses, the validator classes Kaapi encodes, the label sets the manual-evaluation guide teaches. What I am proposing is the conversion step, that they become Guardrails Hub validators, the way Uli's slur list already did. The example config names the validator uli_slur_list. The dataset lives in the product code, an annotation artifact born in research and now running as a gate inside a deployed guardrail system. The whole argument, rendered as a config line.
# Does it hold? Three documented cases
Three documented cases, three tests each: does the framework surface the harm, prescribe a response, and extend past what the original response achieved? I chose these three because they are the only documented social-sector chatbot evaluations with enough published detail to test against: the nutrition out-of-scope case and the caste personal-information case from Baarish's manual-evaluation guide, and the IITM admissions misuse case from Tattle's services page. A framework that cannot reproduce what already happened cannot be trusted on a new deployment. The framework grew out of these same cases, so passing them only shows the minimum.
My position underneath the passes is unchanged from the first analytic memo I wrote after the five assigned readings. Quality over quantity makes definite sense, but the utility of the quality must be scalable. We need a pipeline and a full design approach, not a heroic one-off. The framework is the attempt at the pipeline.
# What's still open
Three things stay open, and I would rather name them than patch over them. Who staffs the annotation labor, and who absorbs its emotional cost. The guide is silent on both, and the sociotechnical paper makes the same point from outside: harm to annotators is named as an evaluation-layer risk in its own right, with annotator compensation left unsolved.
And the one I have not resolved even with myself is disagreement-as-product as it was with the Uli Dateset. This is hard to resolve. As an engineer, product utility is my measure, and here it falls short. But as a dataset that can inform model learning, I understand the immense value qualitative disagreement data may eventually have for model training.