Intent Cannot Be Fully Specified Up Front: Why Specifications for Open-Ended Tasks Cannot Be Completed Before Exploration
Examines why intent cannot be fully specified in advance when generative AI is used for open-ended tasks—how concrete outputs shape judgment, why local feedback cannot be promoted directly to global rules, and how exploration creates new problem dimensions, reorganizes intent, and can ultimately invalidate the overall framework—while also considering the same structure in human communication and joint activity.
- First published
- Last updated
On this page
An asymmetry often appears when generative AI is used for design, writing, research, or complex software development. Before an output exists, users cannot accurately describe what they want or enumerate everything they do not want. Once an output exists, however, they can identify which parts are wrong, which parts are worth keeping, and even criteria they had not previously recognized.
This is not merely a matter of a prompt omitting requirements that already exist. A concrete output reveals the AI’s interpretation of the requirements, the real costs of a proposed solution, and local conflicts. It also participates in shaping subsequent judgment. The same choice may receive opposite evaluations in different local contexts. As outputs and technical understanding accumulate, what changes may be not only the local specification, but also the intent, the problem definition, and the entire solution framework.
AI makes this structure especially visible, but did not create it. Shared understanding, planning, and action between people likewise depend on subsequent interaction and concrete situations. Understanding this phenomenon therefore requires more than asking why expressions omit information. It also requires examining how judgment forms around results, why local evaluations cannot be generalized directly, and how exploration can produce new dimensions of evaluation and new conceptual tools, ultimately changing the problem it initially set out to solve.
Intent cannot be fully specified up front
This article uses the phrase intent cannot be fully specified up front to describe the asymmetry between specification before a result exists and evaluation after it appears. The phrase is not an established technical term with a uniform definition in any one discipline. It synthesizes a set of related problems studied in design research, requirements engineering, preference construction, interactive machine learning, and human–computer interaction.
“Cannot be fully specified” does not mean that nothing can be stated before a task begins, nor that requirements can never become clear. Positive descriptions, negative exclusions, examples, and constraints can all narrow subsequent attempts within the current problem definition. The issue is that they can contain only what the user is able to conceptualize and express at that point. Before concrete results have appeared and technical understanding has accumulated, these expressions cannot be expected to cover local judgments, conflicts, and wholesale redefinitions that will arise only later.
In human–AI collaboration, it is useful to distinguish two concepts:
- Intent comprises a person’s sense of the direction, boundaries, trade-offs, and judgments relevant to a task and its outcome. Some of it may already be stable; other parts can only be identified, clarified, or changed through concrete results and the actual course of the work.
- A specification is an operational expression of that content at a particular moment, such as a prompt, constraint, example, counterexample, test, or acceptance criterion.
A specification can be rigorous or even formally complete within the bounds set by current understanding, but that does not guarantee that it covers every factor that will affect judgment. Conversely, intent is not an invariably stable object already stored in full inside the user’s mind, waiting only for language to reproduce it without loss. Each specification should therefore be understood as an expression of intent at the current stage of understanding, not as a complete copy of intent itself.
Requirements formation is likewise not the transcription of a complete set of demands that already exists. It requires identifying the participants involved in the task, understanding their respective purposes, and translating that content into expressions that can be analyzed, communicated, and implemented. The difficulty arises not only from possible conflicts among different purposes, but also because goals may not yet have formed explicitly, or may have formed to some degree while remaining difficult to express.
This also requires distinguishing two questions: whether the current result conforms to the current specification, and whether the current specification is still sufficient to express the direction worth pursuing. Tests and acceptance criteria can check the former strictly, but they cover only what has been written into the specification. A result may satisfy every existing assertion while the task as a whole still rests on a problem model that omits a critical condition or has ceased to be valid. Conformance to a specification does not prove that the specification adequately expresses intent; the specification itself can become an object of subsequent evaluation.
This incompleteness does not arise from a fixed set of independent gaps. Difficulty conceptualizing a goal, for example, further constrains the ability to express it. New judgments may form only after concrete results appear and local relationships become visible. As requirements and technical understanding continue to change, judgments formed earlier may in turn cease to apply. In practice, these effects intertwine rather than appearing as a sequence of separate categories.
Positive specification is constrained by both conceptualization and expression
Directly describing the desired result can be called positive specification. Such a description assumes that the user can first conceptualize the target with reasonable accuracy and then externalize it in language. When people use AI, that assumption does not always hold. A user may have only a vague direction, scattered judgments, or a handful of references rather than an internal model that can be fully articulated.
An inadequate conception of the final form also impoverishes the means available to express it. This resembles the way people borrow familiar things as analogies when they lack an established concept: “like A, but without these parts of A,” or “somewhere between A and B.” This does not imply that the target itself must be something novel. It means only that when users cannot conceptualize the target precisely, they also lack expressions that can refer to it directly. An analogy can provide direction without thereby becoming a complete specification.
Further clarification can therefore address only part of the difficulty. If a user has already formed a judgment but has not found an adequate way to express it, questions, restatements, or examples may help externalize it. If the judgment itself can form only after concrete solutions are compared, actual consequences are observed, or technical understanding is accumulated, questions do not create those conditions. Asking users to answer every detail in advance merely recasts exploration that has not yet happened as a set of questions they must answer by guessing.
Negative specification is easier to express, but impossible to exhaust
When positive specification stalls, people turn to describing outcomes they do not want. Negative specification does not require users to construct the complete correct answer first; it requires only that they identify errors, directions, or boundaries they can already recognize. It is therefore usually easier to express clearly than a positive description. As long as the current problem definition remains valid, each effective exclusion may narrow the range of subsequent attempts and accelerate movement toward a final form.
Negative specification is nevertheless incomplete as well. The possible errors, combinations, and local conflicts far outnumber the prohibitions a user can anticipate. Many things a user “does not want” become recognizable only after they actually appear in a concrete result. Negative specification is therefore a useful means of approximation, but enumerating exclusions cannot uniquely determine the desired result.
Nor does a negative judgment necessarily exclude only a candidate within a fixed problem space. Sometimes what is rejected is a task decomposition, a solution premise, or a mode of evaluation shared by all the existing candidates. Adding more prohibitions would still move only within the same set of premises. Such a rejection does not further narrow the original range; instead, it requires the problem space to be reopened. Negative specification may therefore either narrow the current solution space or reveal that the solution space itself has been defined inappropriately.
Concrete outputs make evaluation after the fact possible
Once AI has produced a concrete output, the user no longer has to construct a complete answer from nothing and can instead judge an object that already exists. Such judgments are often local: some parts are wrong, some can be retained, and others need another version for comparison. A user may be able to point out exactly what is wrong in one place while remaining unable to describe the correct replacement.
Concrete outputs also expose solution choices, actual costs, and local conflicts that were not anticipated in advance. A choice that appears feasible in the abstract may reveal its relationship to surrounding parts, and the costs it creates, only after it becomes part of the overall result. This newly available information forms the context for the next evaluation, but does not automatically become a complete specification.
A concrete output also externalizes the interpretation that the AI actually adopted in that round of generation. The scope of a goal, the priority among requirements, and unstated premises may remain ambiguous in a prompt, yet appear as observable choices in the result. A generated result is therefore not only an attempted solution but also a concrete interpretation of the current specification. What the user rejects may be not merely a local output, but also the interpretation and premises that produced it.
Concrete outputs, however, are not neutral probes that simply read pre-existing intent. The options, structures, and contrasts they present affect which differences become salient first. They may also turn inclinations with no prior ordering into a concrete trade-off during comparison. A single direction may narrow subsequent exploration; multiple candidates may share the same premises and therefore leave other directions outside the visible range. Evaluation after the fact may consequently identify and express judgments that already existed, but it may also form or change judgments within the context created by the current result and the act of evaluating it. Even when a user can now make a clear choice, this does not establish that a complete and stable set of preferences existed beforehand.
Here, “after the fact” does not refer only to final acceptance after the task is complete. It occurs after every intermediate result. AI produces a version; the user accepts or rejects parts of it; the next version is then generated in light of those evaluations. As long as the current problem definition and overall framework remain valid, such a loop may gradually approach an acceptable result. A single evaluation, however, is not equivalent to a complete specification in advance.
Local evaluations cannot be promoted directly to global rules
Evaluation after the fact has another difficult property: situations that appear identical in an abstract description may receive different judgments in different local contexts. A structure, wording, or implementation technique rejected in one place may be desirable in another. “Do not use X here” therefore cannot be rewritten directly as “never use X.”
Being able to judge a concrete result is not the same as being able to generate a rule in advance that covers all similar results. The concrete object presents its location, purpose, surrounding structure, and actual consequences together, and these conditions jointly support the present judgment. Once the judgment is abstracted into language or a threshold detached from the object, it becomes necessary to state which conditions matter, how they combine, and what exceptions exist. Recognizing whether a local result is appropriate and constructing a generally applicable decision rule are not the same ability. A user can therefore reliably say “this is wrong here” in the current context while being unable to state “what is right everywhere.”
This difference does not necessarily mean that the user is inconsistent. It may depend on the local function, the surrounding content, or the overall combination; the judgment itself may also have changed through iteration. Whatever the reason, a local evaluation cannot automatically become a generally applicable preference once detached from the object it evaluated.
Nor does the inability to promote a local evaluation directly to a global rule mean that evaluation always occurs within a problem space that has already been completely defined and remains fixed. A concrete result may turn a difference absent from the prior specification into a new dimension of evaluation, or change the level at which a problem should be addressed. The question is then no longer only where a particular judgment applies, but also how the current task is being divided and understood.
Open-ended collaboration does not unfold within a fixed problem space
If a task consisted only of locating an unknown target within a problem space whose dimensions and boundaries were already fixed, then positive descriptions, negative exclusions, and candidate evaluations—though incomplete—could in principle still be understood as adding constraints, progressively narrowing the range, and eventually approaching the same pre-existing target. Open-ended collaboration often does not satisfy this premise.
A concrete result may turn properties, relationships, and consequences that were absent from the specification into problems that must now be addressed. A design prototype can expose new contexts of use; a software implementation can reveal new technical constraints; writing and research can force participants to distinguish concepts or levels of a problem that had previously been conflated. What participants gain is not merely new information about existing options, but also new dimensions of evaluation, new ways of categorizing the problem, and the conceptual tools needed to express intent. Some judgments become available as distinct objects of thought and expression only after these distinctions have formed.
Intent therefore does not always change by gradually converging along a fixed direction. A new specification may do more than add constraints to the old one. New problem dimensions may reinterpret earlier trade-offs or reveal previously unseen dependencies, priorities, or conflicts among requirements that had been listed separately. They may even mean that the collaboration no longer faces the task as it was originally framed. Intent is reorganized through this expansion rather than merely becoming more precise within an unchanging space.
These dimensions and concepts emerge through a particular trajectory of candidates, comparisons, and evaluations, so the resulting intent is path-dependent. The latest specification can record current conclusions, but it may not preserve which candidates made a distinction important, which earlier categories were abandoned, or why a judgment applies only within a particular scope. The resulting intent depends not only on which candidates were seen, but also on the way participants learned to understand the problem during exploration. It therefore cannot always be compressed without loss into the latest specification detached from the history of collaboration.
A changing problem space does not necessarily require the current solution to be abandoned. The original framework may still absorb the new dimensions. But when the problem definition, specification, and intent that emerge later can no longer be accommodated by that framework, a local conflict may become the first place where failure of the whole framework becomes visible.
From local repair to redefining the entire process
Local iteration does not always continue converging within the same framework. An almost complete product may undergo a chain of changes when an attempt is made to solve one small problem, bringing most of the preceding work—and the entire process of understanding the requirements, choosing a solution, and carrying it forward—back into question. At the same time, it is precisely the requirements clarified and the technical understanding accumulated during that process that allow the whole process, even if its outcome is discarded, to become material for redefining the problem and beginning another cycle.
How a local problem exposes failure of the overall framework
Within the current problem definition and solution framework, continued iteration advances the product while also changing the basis on which the product is evaluated. Each output makes previously unobservable trade-offs, conflicts, and consequences concrete. Evaluations of those outputs supplement, revise, or even reorganize the existing specification and may further shape intent itself. At the same time, implementation, investigation, and failure produce new technical understanding, gradually revealing constraints, costs, and alternative paths that were previously invisible.
Consequently, by the time a product is nearly complete within its current framework, the task’s specification, intent, and technical context may differ from what they were when that framework was selected. “Nearly complete” means only that little work remains relative to the specification that has gradually formed and been incorporated into the current solution. It does not establish that the framework remains appropriate for the problem as it now stands. A framework embodies an earlier problem definition, boundary, and set of core trade-offs, but does not automatically change with every new understanding. The framework and the problem may therefore drift out of alignment over the course of iteration.
A seemingly minor problem is not the root cause of this misalignment. It may instead be the first conflict that the existing structure cannot absorb locally. Modifying it affects adjacent parts, which in turn depend on earlier structures, assumptions, and trade-offs. As the scope of modification expands, scrutiny of earlier work moves progressively upward, until what is destabilized is no longer the original local detail but the framework supporting most of the preceding work.
Such late-stage abandonment cannot be attributed simply to a poor initial framework. A framework may have been reasonable and coherent, and even well implemented, given the specification, intent, and technical understanding available at the time. Its later failure does not invalidate its earlier rationality. What fails is the correspondence between the framework and the problem that has formed through iteration. The understanding of the requirements, the problem model, and the technical premises on which the framework was selected can no longer support the current specification and intent. This is failure at the level of the foundation, not a defect in one implementation inside the framework.
A better initial design still has value. It can reduce unnecessary coupling, establish boundaries that make replacement possible, and lower the cost of wholesale abandonment. But it cannot contain in advance the requirement judgments and technical understanding that can be acquired only through actual iteration. It therefore cannot guarantee that the initial framework will remain valid once that understanding appears.
Nor is this chain failure merely excessive coupling in the implementation, and it cannot necessarily be resolved through refactoring. Refactoring normally preserves the current problem definition and principal behavior while reorganizing the implementation. Here, the requirements, the product form, the understanding of the technical context, the architectural choice, and the core trade-offs all need to be reconsidered. All the work completed so far may have been produced within a framework that no longer holds.
Once the framework is reconsidered, the next step may adopt a solution that has never been used before, or return to one considered and rejected earlier. Redefining the requirements may even reveal that a task previously addressed with a complex system can be eliminated, combined with something else, or accomplished with an extremely simple solution.
How the entire process becomes input to the next cycle
The requirement judgments and technical understanding formed through iteration do not reside only in the final product or the latest specification; they are distributed across the entire cycle. Beginning with the initial vague description, each act of generation and evaluation makes some part of the requirements clearer. Each implementation, conflict, and failure supplies technical context that was previously missing. The actual costs of different solutions, the dependencies among local choices, and the limitations introduced by early assumptions can often be seen only in relation to the path that was actually taken.
What must be reviewed, therefore, is no longer merely the final result but the whole cycle: how the problem was initially defined, which requirements could not yet be expressed, which unverified assumptions entered the solution, why the current path was selected, and which local feedback was generalized incorrectly. Only after traversing the entire process may it become apparent that the problem lies not merely in defects in the current product, but in the overall way this cycle understood the requirements, selected a solution, and continued patching it.
This realization changes the nature of all the work that came before. The old implementation may not be reusable, and the current product may be discarded entirely. Yet earlier requirement statements, candidate results, local evaluations, failed modifications, technical investigations, and abandoned directions together constitute the material needed to redefine the problem. The first cycle is no longer merely a failure that produced no deliverable; it is a complete exploration that made the new definition possible.
The next cycle is therefore neither a continuation of the old solution nor a return to the original starting point with everything erased. It takes the whole previous process as input, redefining the problem and selecting a path with clearer requirements, richer technical context, and a reopened solution space. The problem it addresses may no longer be the task as understood at the beginning of the previous cycle. The results and process of this new cycle may later be reevaluated as a whole in the same way.
The inability to fully specify intent up front is not unique to AI
This problem often becomes especially visible when generative AI is used, but it is not necessarily caused by AI. Nor can one person copy the entire contents of their mind into another’s simply by uttering a sentence. In open-ended joint activity, the purpose, division of labor, and course of action are not always settled before communication begins. To determine whether this is an inherent property of communication, four different objects must first be distinguished:
- Private intent is what one party currently seeks to express or accomplish. It may include both judgments that are already stable and matters that remain undecided.
- An expression is an external object produced from private intent through speech, documents, examples, or plans. Another party can perceive it, but it is not identical to the private intent.
- Shared understanding is the participants’ current agreement about what has just been expressed and what should guide the next action. It does not arise automatically once an expression has been issued.
- A joint plan is how the participants intend to coordinate subsequent action. It is influenced by each participant’s intent and continues to unfold through joint decisions and actual activity.
An expression may not carry private intent in full, and its recipient may form a different interpretation. Even after the participants have established a shared understanding adequate for the moment, subsequent action may still alter each participant’s intent and their joint plan. The problem in human communication is therefore not only that information may be lost in transmission. Shared understanding itself can be established only through interaction, and some of the intent required for action has not yet formed.
Shared understanding is not the result of a single transmission
Clark and Wilkes-Gibbs’s study of how people establish reference observed how two participants used dialogue to identify and arrange complex figures. The speaker generally did not begin with an exhaustive description, but proposed a relatively simple referring expression. The listener could accept it, ask a question, supplement it, or correct it, and the pair would revise the expression iteratively until they reached a reference they both accepted for the moment. An expression here is not a finished product independently created by the speaker and then handed to the listener for decoding. It is a contribution that the two participants must complete together.
Clark later used grounding to describe this process of establishing understanding together. Participants seek to add something to their common ground, but only to a degree sufficient for current purposes. The normal stopping condition for communication is therefore not proof that the participants possess identical internal understandings, but enough evidence for the current joint activity to continue. What has been achieved is a local and temporary closure. An understanding that was previously adequate may be reopened when the task, object, or risk changes.
This common ground also depends on the participants’ interaction history. Brennan and Clark’s experiments on “conceptual pacts” found that interlocutors develop partner-specific terms for repeatedly discussed objects and continue using them in later interactions. An expression that has become clear within one collaborative relationship does not necessarily carry the same meaning for a different participant. Shared understanding between people has a scope of application and cannot be detached from its participants and formation process and promoted directly to a global definition.
The interaction history preserves more than the term ultimately retained. Expressions that participants previously accepted, revised, or abandoned establish contrasts, reasons, and scope for the current wording. The last utterance or latest document can record the current conclusion, but does not automatically carry these conditions of formation. Current shared understanding and the joint intent built upon it therefore cannot always be compressed without loss into the latest specification. Part of that shared meaning still depends on the candidates, rejections, and repairs the participants experienced together.
Repair is part of communication
Incomplete understanding is not an occasional event confined to complex projects. Dingemanse and colleagues analyzed natural conversation in twelve languages from eight language families across five continents. Every sample exhibited a similar structure of other-initiated repair: the listener signals a problem with the previous turn, and the speaker repairs it through repetition, clarification, or confirmation. Such repairs appeared in the samples, on average, about once every 1.4 minutes.
If communication is understood as the transmission of complete information, repair can only be a remedy for a failed transmission. Dingemanse and Enfield’s review of the relevant research offers the opposite interpretation: interactive repair is part of the infrastructure that allows human language to remain complex, flexible, and resilient. Because participants know they can ask questions and make corrections if an expression causes a problem, they need not enumerate every piece of background, every condition, and every exception in each utterance. They can normally begin with a lower-cost expression and add the necessary information only when an actual discrepancy appears.
From this perspective, incompleteness is not always a defect in communication that should be eliminated. Expression and repair together form a mechanism for distributing cost. The first expression establishes enough direction to begin; subsequent interaction discovers which omissions now interfere with understanding and adds content selectively. Human communication normally works through this mechanism, not through a message that can be complete on its own without feedback.
Joint intent may also form through collaboration
Interactive repair can still occur when private intent already exists in full: the speaker knows what they mean, but the listener has not understood it. Open-ended collaboration presents a stronger case. The participants themselves may have only partial plans and must make joint decisions before they can determine what they intend to do next.
Grosz and Hunsberger’s formal study of the dynamics of collaborative intent begins with resource-bounded participants and a continuously changing environment. It allows a group to form an incomplete joint plan in which individual intentions remain underspecified. Subsequent group decisions extend the joint plan and require each participant to update their individual intentions accordingly. The study does not describe one party conveying a complete plan to another, but multiple partial plans gradually forming a jointly executable structure through coordination.
Nor can what is formed together be attributed straightforwardly to either participant alone. Carassa and Colombetti distinguish the private meaning a speaker seeks to convey from the shared meaning that the parties establish by accepting an interpretation. The two can coincide, but they are not the same object. In organizations, participants from different specialties may also understand the same product and problem through different work settings. Research on engineers, technicians, and assemblers on a production floor found that they established common ground while resolving specific misunderstandings and thereby changed one another’s understanding of the product and the production process. In such cases, communication does not merely transmit existing understanding; it participates in generating new shared understanding.
People “aligning on intent” therefore need not end with identical copies of intent in their respective minds. A more realistic outcome is that the parties retain different knowledge, judgments, and local purposes while establishing enough shared understanding, commitment, and partial planning to support the next action. New decisions and consequences will continue to modify this temporary coordination.
A plan cannot replace continued judgment in context
Even when participants have formed an explicit plan, that plan cannot determine subsequent action in full. Suchman treats plans as resources for situated action rather than complete scripts for action. Plans can remain concise precisely because they do not represent all the concrete conditions of real activity. Their prescriptive force necessarily remains vague, and actors must continue to judge the environment and consequences they actually encounter.
Schmidt and Bannon use articulation work to describe how participants in real collaboration manage task dependencies, unforeseen failures, differences in professional judgment, and local conditions. Real work systems are open, and formal procedures cannot guarantee that every contingency will be covered in advance. Collaboration requires participants to continue creating local and temporary arrangements so that work can proceed amid incomplete knowledge and inconsistent conditions. Their CSCW study treats articulation work as inseparable from collaboration itself, not as an additional cost that a more detailed plan could eventually eliminate.
This has the same structure as the inability to fully specify intent up front. An advance expression provides a starting point for action; the actual situation produces information that did not previously exist; the participants then revise their understanding, plans, and purposes in response. Requiring a plan to contain all these judgments before action begins is still requiring the result to precede the situation that produces it.
AI exposes and amplifies this structure
AI did not create the preceding problems, but common prompt-based interaction weakens or conceals mechanisms that human communication normally uses to handle them. People can use pauses, hesitation, questions, restatements, and observation of action to keep judging whether their shared understanding is adequate. They are also constrained by shared commitments and relationships of responsibility. Current generative AI, by contrast, may immediately produce a formally complete and confidently phrased result before shared understanding has been established. Its fluency can easily be mistaken for evidence that understanding exists, even though the model’s interpretation of the instruction, its implicit assumptions, and its uncertainty have not thereby become genuine common ground between the parties.
At the same time, AI can expand one interpretation into concrete text, design, or implementation at very low cost. Differences that remain implicit in human collaboration therefore become visible results more quickly: only when the user sees the result do they discover that the interpretation is wrong, after which the model generates another version from the local rejection. AI makes this loop more intensive. It also increases the risk of generalizing local feedback incorrectly and rapidly accumulating work within the wrong framework.
AI is therefore better understood as an amplifier of an existing structure of communication. Natural language already depends on contextual inference and interactive repair. Shared understanding can already be established only provisionally, and joint plans in open-ended collaboration already change through action. The problem with the prompt paradigm is that it mistakes one expression within these processes for a specification that can be complete independently of them.
This does not mean that every form of communication is equally resistant to specification. Communicating an already determined fact, exchanging data under a closed protocol, or defining formal requirements within a fixed problem model may all yield a specification sufficiently complete relative to that model. What pervades natural-language communication is the dependence of expression on context and common ground. The stronger phenomena of intent formation and wholesale redefinition occur chiefly in open-ended joint activities where the space of possible results has not been explored, participants have limited knowledge, and the environment returns new information. The inability to fully specify intent up front is not an undifferentiated property of every utterance, but an irreducible feature of the relationship between communication and action when these conditions occur together.
Related research
The concepts discussed in this article have counterparts in requirements engineering, design research, pragmatics, psycholinguistics, computer-supported cooperative work, preference construction, interactive machine learning, and human–computer interaction. The literature below is organized by the concepts for which it is used here.
A specification expresses intent at the current stage of understanding
Bashar Nuseibeh and Steve Easterbrook. Requirements Engineering: A Roadmap. Proceedings of the Conference on the Future of Software Engineering, 2000.
This roadmap surveys the research foundations, working contexts, and major activities of requirements engineering from the standpoint of the real-world purposes a software system is expected to achieve. It treats requirements elicitation, modeling and analysis, communication and negotiation, and requirements evolution as interleaved activities spanning the system lifecycle. It also discusses scenarios, prototyping, goal modeling, formal specification, requirements validation, conflict management, traceability, and change management. The paper emphasizes the multidisciplinary character of requirements engineering: beyond software and systems engineering, it draws on cognitive science, the social sciences, linguistics, and philosophy to understand individual knowledge, organizational relationships, and constraints in the real world.
The roadmap preserves an ongoing process of translation between stakeholder purposes and specifications that can be analyzed, communicated, and implemented. Its authors deliberately use “elicitation” rather than “capture” to avoid suggesting that requirements already exist in complete form and can be collected unchanged if only the right questions are asked. Goals may not yet be explicit or may be difficult to articulate, while models and prototypes can in turn elicit new requirements. A specification can therefore be highly precise within a given formalism and still require validation against reality, as well as continued evolution with its participants and environment. Formal completeness describes how the current expression is organized internally; by itself, it cannot establish that the expression has copied the intent behind it in full.
Yoonsu Kim, Kihoon Son, Seoyoung Kim, Brandon Chin, and Juho Kim. IntentFlow: Investigating Fluid Dynamics of Intent Communication in Generative AI. Proceedings of the 2026 Designing Interactive Systems Conference, 2026, pp. 3804–3837.
This study examines how intent is expressed, explored, preserved, and resynchronized in interactions with generative AI. The authors first systematically reviewed forty-six papers published at ACM human–computer interaction conferences from 2022 through 2025 and organized intent communication into four activities: expression, exploration, management, and synchronization. They then implemented the IntentFlow research prototype and used a within-subject comparison with twelve participants to analyze the sequences of actions taken while completing open-ended writing tasks in a conventional chat interface and in the prototype. The study found that users did not communicate their requirements only once at the outset, but moved repeatedly among generation, inspection, revision, and organization. An interface that supported explicit revision and preservation of intent also made users less dependent on successive correction and allowed them to organize the goal structure they had formed over time.
This cyclic behavior suggests that the current input serves more to establish a temporary state for the next generation than to deliver a complete set of requirements once and for all. Generated content may cause users to add constraints, alter relationships among preferences, or return to an earlier direction and reopen exploration, so the operational expression within a single task is continually rewritten. The paper, however, calls higher-level and relatively stable objectives “goals,” while reserving “intent” for the changing strategies, preferences, and constraints involved in pursuing them. Its use of “intent” is therefore narrower than this article’s. It directly demonstrates the fluidity of lower-level intent, but cannot by itself establish that higher-level goals also necessarily remain open. Its four activities are likewise a framework for analyzing generative-AI interfaces, not an exhaustive classification of the sources of incomplete intent.
Negative specification is easier to express, but impossible to exhaust
Yoonsu Kim, Jueon Lee, Seoyoung Kim, Jaehyuk Park, and Juho Kim. Understanding Users’ Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level. Proceedings of the 29th International Conference on Intelligent User Interfaces, 2024, pp. 385–404.
This study collected 307 conversations in which 107 users employed ChatGPT for real tasks and analyzed 511 instances of dissatisfaction within them. The authors classified the types of dissatisfaction and the tactics users adopted in subsequent interaction. Problems in understanding user intent were the most common source of dissatisfaction. Faced with an unsatisfactory result, users repeated the original prompt, made their original intent more specific, identified and corrected errors in the result, or changed the task itself. The paper also examines the effect of domain knowledge: how much users knew about the subject influenced whether they could identify a problem, what kind of correction they could provide, and whether the dissatisfaction was ultimately resolved. Despite these tactics, only a portion of the observed dissatisfaction was resolved in the ensuing dialogue.
Once a result exists, the previously abstract sense that it “does not meet expectations” acquires a concrete object. Users can then turn it into actionable information such as “this is what was misunderstood,” “this fact is wrong,” or “this is how the task should change.” These behaviors show that making intent more specific and pointing out errors are common ways to proceed when a positive description is insufficient. The large number of problems that remained unresolved after multiple corrections also shows that such exclusions do not automatically add up to a complete specification. The study observes how users handle dissatisfaction that has already occurred, however, and does not directly compare the expressive difficulty of positive and negative specification. It provides empirical evidence about correction after a result and its incompleteness, while the claim that negative specification is usually easier to express still requires a separate argument.
Concrete outputs make evaluation after the fact possible
Donald A. Schön. Designing as Reflective Conversation with the Materials of a Design Situation. Research in Engineering Design, 1992.
Schön considers how artificial intelligence might represent, simulate, or assist the “knowing-in-action” of architectural designers and distinguishes four objectives: functional equivalence, phenomenological equivalence, design assistance, and a research environment. Through teacher–student dialogue in design studios, controlled design exercises, and “design games,” the paper analyzes how designers actually work. Designers place lines and shapes within “design worlds” of their own construction, see new figures and properties in what they have made, and then test those perceptions through the next move. The problem, intent, and evaluation also change with action and its unintended consequences. On this basis, the article questions whether design activity can be replicated through an exhaustively predefined set of symbolic rules and argues that computing systems should be developed more as design assistants and environments for studying design cognition.
The “seeing–moving–seeing” process shows that a sketch or prototype is not merely the passive output of an already complete intent. Action first creates a situation in the material that did not previously exist; only then can the designer see proportions, conflicts, possibilities, and unintended consequences and form a judgment that can guide the next move. Being able to see these properties in concrete material also does not mean that the designer could have translated them in advance into decision rules covering every case. The concrete result simultaneously changes the object being evaluated and the design situation as understood by the evaluator. Some criteria are not prewritten and merely omitted, but acquire form only through the perceptible situation. Evaluation after the fact contains more information than a description in advance not simply because the wording is now easier, but because neither the object available for evaluation nor the understanding it evokes existed before.
Sarah Lichtenstein and Paul Slovic. The Construction of Preference: An Overview. In The Construction of Preference, 2006.
This introduction to the volume reviews the central questions and evidence in research on the construction of preference. The authors do not deny the existence of stable preferences. Instead, they distinguish choices for which existing preferences provide a direct answer from situations in which a response must be constructed during judgment: the elements of the decision may be unfamiliar, several existing preferences may conflict without an established weighting among them, or a person may have a clear positive or negative feeling but struggle to convert it into the numerical response required by the task. Using the chapters in the volume as a guide, the introduction then examines preference reversals; descriptions of options and questions; option sets and decision contexts; emotion and reasoning; and the way predictions of future experience alter choice. It shows how different modes of elicitation can produce systematically different responses and make preferences appear mutable, inconsistent, and context-dependent.
A concrete set of candidates turns an abstract direction into an actual trade-off: which attributes occur together, how strongly they conflict, and what form the judgment must take are established only at that point. People may possess several stable inclinations in advance without knowing their relative weight in the current combination, and may notice an unconsidered attribute only when they see the options. Evaluation after the fact therefore neither creates preference from nothing nor always reads a complete standard that was stored in advance. It often completes a judgment under conditions jointly provided by prior inclinations, the current object, and the mode of evaluation. The candidate set is thus not merely a carrier on which preference is displayed; it participates in constituting the context in which preference is expressed and formed. Making a clear choice among current options does not directly establish that a complete ordering exists independently of them. An inability to state every choice rule in advance can therefore coexist with a clear evaluation once the results are visible.
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems, 2017.
This study proposes a method for training reinforcement-learning agents from human comparisons of behavior segments. The system selects pairs of one- to two-second trajectory segments generated by the current policy and asks an evaluator to choose the better one, mark them as equally good, or indicate that the choice cannot be made. These comparisons continually train a reward model, and the policy then optimizes the reward predicted by that model. The authors test the loop in simulated locomotion and Atari games. With human feedback on less than one percent of the agent’s total interactions, the system learned a range of behaviors that were difficult to encode in a reward function directly, including a backflip for which no ready-made objective function existed. The experiments also show that freezing the reward model without continued feedback can allow the policy to discover anomalous behavior that humans do not endorse.
The evaluator is not asked “what reward should every possible behavior receive?” but “which of these two behavior segments is better?” A behavior segment presents motion, context, and concrete consequences together, turning a difference that is hard to formalize into a question that can be compared. The fact that longer segments were sometimes more useful also indicates that evaluation depends on visible local context. Continuing to generate new behaviors and request judgments is equally important, because old comparisons cannot cover in advance the behaviors a policy will later discover. The experiment does not recover a complete intent from these labels. It shows instead that relative judgments about concrete results can supply information absent from an advance reward specification and can serve as input to the next generation.
Conformance to the current specification does not prove that the specification itself is adequate
Shreya Shankar, J. D. Zamfirescu-Pereira, Björn Hartmann, Aditya G. Parameswaran, and Ian Arawjo. Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 2024.
This paper studies how developers create automated evaluators for particular LLM applications and how those evaluators can be made to correspond to developers’ judgments of actual outputs. The authors designed the EvalGen workflow, in which developers first inspect a set of outputs and make binary judgments, then propose natural-language criteria that an LLM converts into executable assertions. The system displays discrepancies between those assertions and the human judgments so that developers can revise the criteria or reconsider their original evaluations. Nine participants with experience developing LLM applications used the workflow with each of two generation pipelines. The study found that participants not only added previously omitted criteria, but also reinterpreted existing criteria after seeing new types of failure and even changed earlier judgments of outputs. The authors call this coevolution of criteria and judgments “criteria drift.”
The evaluation criteria here are not completed first and then held fixed for every output. An executable assertion can check one criterion consistently without establishing that the criterion covers the judgments the developer actually applies. A discrepancy between the two may require changing not the output, but the criterion used to evaluate it. Developers first need to judge concrete results before they can see which distinctions deserve to become criteria. Applying a new criterion to more outputs may then reveal overgeneralization, conflicts, or gaps in coverage, forcing them to revise both the criterion and their earlier judgments. Concrete results are therefore not only objects of evaluation; they also participate in forming the scale by which they are evaluated. The study includes only nine experienced participants, two pipelines, and short tasks, and does not cover long-term evolution after deployment. It is sufficient to show that criteria can form and drift during evaluation, but not to estimate how prevalent this phenomenon is across all open-ended tasks.
Local evaluations cannot be promoted directly to global rules
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell, and Anca Dragan. Inverse Reward Design. Advances in Neural Information Processing Systems, 2017.
This paper addresses reward misspecification in reinforcement learning. A designer will often write a convenient proxy reward in a training environment, but an agent that maximizes it literally in a test environment containing new features may cause side effects or exploit loopholes in the reward. Inverse reward design does not treat the proxy reward as the true objective. It treats the proxy instead as evidence of a reward chosen by a designer acting approximately rationally within the known training environment, uses Bayesian inference to recover possible true rewards, and preserves uncertainty during planning. The authors also propose a method for approximating this inference through inverse reinforcement learning. In small environments containing unseen dangers such as lava and exploitable features, they demonstrate that inference combined with risk aversion can reduce catastrophic behavior compared with executing the proxy reward directly.
A distinction absent from the training environment cannot affect which proxy reward the designer writes there. Multiple true objectives may consequently support the same specification in that environment. If a system later interprets the original specification’s silence about that distinction as an unconditional preference in a new environment, the system has added its own extrapolation rather than recovered a judgment the designer made. Local evaluation has a similar boundary. It can reliably distinguish conflicts that actually appear in the current result, but does not simultaneously answer every trade-off in contexts that have not appeared. Inverse reward design is not a general model of everyday feedback, but it demonstrates precisely how information sufficient in its original context can acquire meanings it never had once it is treated as a global objective.
J. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. CHI, 2023.
Using a no-code chatbot design tool, this study observes how ten participants with almost no prompt-design experience modified a system prompt so that a cooking assistant would exhibit the instructional characteristics of a specified chef. The tool supported both immediate test conversations and systematic replay tests with error marking; in a time-limited think-aloud task, participants relied mostly on the former. They generally modified the prompt after one or two conversations, treated one success as evidence that the problem was solved, rarely used systematic testing, and tended to infer how the language model would interpret an instruction from their experience of human communication. As a result, a modification made for the current output might not generalize to other inputs and could even reintroduce a problem that had disappeared earlier. The authors frame this small study, whose participants mostly had technical backgrounds, as a formative investigation and use it to discuss how prompting tools and instruction might help users develop more accurate models and testing practices.
Participants could make meaningful criticisms of the output in front of them. The difficulty arose in the next step: did the criticism apply only to the current response, to a class of inputs, or to the behavior of the chatbot as a whole? A change working in the current conversation did not provide evidence that it would continue to hold in other contexts; later regressions made this mismatch of scope directly visible. The problem cannot be reduced to users being unable to find errors. Local errors were in fact easy to see, while abstracting a stable rule from one observation required new samples, comparisons, and tests. Evaluation after the fact thus provided new information, but the scope in which that information remained valid still had to be established through subsequent results.
Silviu Pitis, Ziang Xiao, Nicolas Le Roux, and Alessandro Sordoni. Improving Context-Aware Preference Modeling for Language Models. Advances in Neural Information Processing Systems 37, 2024.
This study addresses context dependence in language-model preference data. Responses to the same prompt do not necessarily have a fixed ranking independent of the user’s purpose, yet conventional preference models commonly compress judgments made under different conditions into a single reward function. The authors treat the prompt as a partial description of user intent and propose providing a preference model with additional context so that it can distinguish evaluations made for different users or task conditions. They construct a synthetic dataset containing preference reversals, in which the ordering of the same pair of responses switches when the context changes, and use it to compare how standard and context-aware models learn and generalize conditional preferences.
Preference reversal turns the scope of local evaluation into a testable structure. If two judgments depend on different contexts, combining them into a global ranking after stripping away those conditions does more than lose detail; it may produce the opposite conclusion. One evaluation can establish that a response is more appropriate under the current conditions without specifying which response should be chosen when those conditions change. The experiments use mostly synthetic reversal data and reward models, and treat intent as a latent variable that is inferred but relatively fixed. The study therefore supports the context dependence of evaluation and the failure of unconditional extrapolation, but cannot explain how intent itself forms or changes through interaction.
The problem space expands and is reorganized through exploration
Horst W. J. Rittel and Melvin M. Webber. Dilemmas in a General Theory of Planning. Policy Sciences, 1973.
Starting from the value conflicts faced in social policy and planning, Rittel and Webber criticize the treatment of open social problems as “tame problems” that can first be defined and then solved by scientific or engineering methods. The paper characterizes “wicked problems” through ten properties: they have no definitive formulation or clear stopping rule; solutions can be judged better or worse rather than simply true or false; consequences cannot be tested immediately or conclusively; possible solutions cannot be enumerated; and every intervention carries real-world cost. Such problems are also unique to their situations, may be symptoms of other problems, and admit multiple explanations, while the explanation adopted directly determines which solution is pursued. The authors use these properties to show that, in planning involving plural interests and normative judgment, problem definition, knowledge production, value trade-offs, and political responsibility cannot be compressed into a problem-solving procedure determined entirely in advance.
When a problem has no final formulation independent of its solutions, continually addressing one local conflict also continually selects what the “real problem” is. The consequences of a new solution may do more than invalidate an implementation; they may reveal that the prior causal explanation, value boundary, or level of analysis is itself untenable. Continuing to patch at that point perpetuates a definition that has already failed rather than approaching the same fixed target. This does not make the earlier attempts useless. Their consequences and the understanding they produced are precisely what allow the problem to be reformulated and another path to begin. Open-ended tasks in AI collaboration are not automatically equivalent to wicked problems in social planning. But whenever the problem formulation similarly depends on knowledge that appears only during attempts to solve it, the same movement from a complete cycle of exploration into redefinition can occur.
Rebecca Fiebrink, Perry R. Cook, and Daniel Trueman. Human Model Evaluation in Interactive Supervised Learning. CHI, 2011.
Through three studies of the Wekinator interactive machine-learning tool, this paper examines how musicians and other creators evaluate supervised-learning models they train themselves. The studies cover seven composers designing instruments over ten weeks, twenty-one students completing course projects, and one professional cellist building a gesture classifier. The authors found that users frequently evaluated correctness, the severity and location of errors, trustworthiness, complexity, unexpected behavior, and practical usability by directly performing with or manipulating a model. Cross-validation scores did not consistently represent these judgments and sometimes correlated negatively with subjective quality when the training data expressed the wrong concept. Evaluation also helped users learn how to collect data, what a model could learn, and how the technology was constrained in a real performance environment. It could change their training method, their use of the system, or even the learning problem itself.
Iteration in these studies did more than improve model parameters. Model behavior also reshaped the user’s understanding of the requirements and the technical context. Direct use might reveal that the concept represented by the training examples was wrong from the outset, or that behavior required by the original plan was infeasible under the current form of interaction. The data representation, mode of use, or task definition then had to change. At that point, the object of evaluation had expanded from a model to be repaired into the entire setup that produced it, while all the prior training and failures became evidence for choosing again. This provides an observable local case in which the complete process, having expanded understanding, becomes material for the next definition of the problem rather than merely a sequence of intermediate versions on the way to a predetermined result.
J. D. Zamfirescu-Pereira, Eunice Jun, Michael Terry, Qian Yang, and Björn Hartmann. Beyond Code Generation: LLM-supported Exploration of the Program Design Space. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, Article 153, pp. 1–17.
This paper extends LLM programming assistance from generating a predetermined piece of code to exploring the program design space. The authors developed the Pail prototype, which lets users examine different problem formulations and solutions side by side, generate and compare executable program sketches, and record requirements and design decisions that emerge during exploration. They then observed eleven participants using p5.js for an open-ended prototyping task and analyzed how the participants moved between problem formulations and solutions. Low-cost, disposable program sketches allowed participants to try substantially different directions quickly. Running and comparing those sketches both created new implementation choices and exposed requirements that the original problem formulation could not accommodate, leading participants back to problem discovery and problem definition rather than continuing to refine code along a single path.
When an implementation is used to understand a problem rather than merely complete it, rejecting that implementation can be more consequential than simply “trying a different implementation.” Programs, changing requirements, trade-offs, and failures accumulated along one path jointly reveal what the original definition omitted and provide material for reopening the design space. Discarded branches can thereby continue to participate in the formation of the next problem. This is a generative-AI-era example of returning from solution exploration to problem definition. The study, however, uses short, small-scale p5.js prototypes. It neither observes nearly complete real products nor tracks cascading failures of an architectural foundation over long-term development. It supports the coevolution of problem and solution, but cannot substitute for this article’s stronger account of an entire framework being overturned and a new cycle of definition beginning.