The method came before the AI: What building an auditable research tool taught me about AI literacy

Post by Michael J. Henderson.

Picture of a research process when a hand inserting a new step
Featured image. The method should survive the model: a research pathway can remain stable while particular AI components are inspected, replaced and governed by the researcher.

When academics first begin exploring generative AI, one of the most understandable questions is: Which AI should I use? ChatGPT? Claude? Gemini? A locally hosted model? Something supplied by the university?

I increasingly think this may be the wrong place to start.

My own route into using generative AI in research began with a method, not with an AI system. The method was designed to show how adults make significant learning choices: what influences those choices, how different factors interact, and how choices develop over time. Only later did I begin asking whether AI might help bring parts of that interpretive process closer to the research interview itself.

That sequence has become important to me.

It means that the question I now ask is less “What can this AI do?” and more “What, within this research method, should I allow an AI to do?”

Those questions sound similar, but they lead in different directions.

The first can encourage us to organise research around the capabilities of whatever product happens to be in front of us. The second requires us to identify the parts of a research process in which AI might be useful, the limits we want to place on it, what evidence it can see, what happens to its output, and which decisions must remain with the researcher.

This has also changed what I mean by AI literacy.

Lynette Pretorius’s work on AI literacy has argued that explicit modelling can help research students understand generative AI as a learning tool, rather than leaving them to discover its possibilities and limitations unaided. My own experience has pushed me to add a further layer.

For researchers, AI literacy cannot simply mean becoming competent at prompting a particular system. Models, interfaces and platforms are already changing too quickly for that to be sufficient. I now think it also requires enough understanding of the architecture surrounding AI to know where it sits in a research process, what authority it has been given, and how its contribution can be inspected, challenged or rejected.

You do not need to become a software engineer to ask those questions.

But you do need to know that there is an architecture underneath the chat box.

The method came before the AI

The Learning Choice Tracker (LCT) did not begin as an AI project.

It emerged from my doctoral research into adult learning choice-making. In its original form, it was a manually constructed interpretive representation developed from qualitative interview material—and even that form evolved over an extended period of design and testing. The purpose was to make visible something that could otherwise remain difficult to hold together: the combination of personal, relational, educational and structural influences contributing to significant learning choices.

The difficulty was temporal as much as analytical.

A conventional qualitative process can produce rich interpretation, but that interpretation is often constructed well after the interview has finished. By the time a visual or conceptual representation is developed, the opportunity to explore the context directly with the participant may have gone.

That disparity existed long before generative AI became part of the discussion and therefore raised a methodological question:

Could the interpretive work be brought closer to the interview without reducing its quality or transferring too much cognitive load to the researcher?

An earlier version of the LCT, documented in a 2023 AARE conference presentation, explored a more immediate workflow using transcription, forms, spreadsheets and visualisation tools. That shift moved the representation closer to the interview, but it also revealed a practical problem. Making a qualitative method faster is not particularly useful if the researcher becomes so occupied with operating the machinery that their attention is pulled away from the participant.

Generative AI changed what appeared technically possible, but it did not change the underlying methodological problem.

The current AI-augmented LCT therefore does not ask an AI to “analyse the interview” and deliver a finding. AI components are assigned bounded tasks: they can propose contributory aspects from transcript evidence and suggest how related aspects might form significant choice zones. The proposals remain linked to evidence and subject to researcher review.

That separation between an AI-proposed interpretation and a researcher-approved interpretation is built into the research instrument. That methodological clarity showed me that methodological principles can be translated into technical boundaries: architecture can determine not only what an AI can do, but what it is permitted to do.

For me, that has become a more useful starting point for research with generative AI than the search for the “best” model.

The method should survive the model.

There is an architecture underneath the chat box

When AI and the LCT “met”, the method’s purpose did not change. AI offered the prospect of greater speed and accuracy, with the potential for improved participant engagement. It also required an architecture capable of protecting the method.

“Architecture” can sound like a layer of complexity best left to software engineers. I use the term more simply here. It describes how the parts of a research process are arranged, what passes between them, and where decisions are made. A researcher does not need to understand every technical detail, but does need to be able to follow the movement of evidence and interpretation through the process.

In the current LCT, that movement can be represented as:

Interview → evidence → AI-proposed aspects → AI-proposed zones → researcher review → approved interpretation → representation

Expressed in more familiar research language, this becomes:

evidence → provisional interpretation → checking → human judgement → accountable representation

Picture of project architecture showing the human being responsible for the final decision
Figure 1. AI proposes; the researcher decides. Interview evidence is retained as AI proposes contributory aspects and choice zones. Researcher review can reject, revise or rerun those proposals; only an approved interpretation proceeds to the visual representation.

Each step marks a methodological decision. The transcript provides the evidence available to the AI. The AI can propose aspects that may have contributed to a learning choice and suggest how related aspects might form a significant choice zone. Those proposals retain links to the evidence from which they were developed. They can then be examined, altered or rejected by the researcher. Only an approved interpretation becomes authoritative for what follows.

Separating generation from approval makes the boundary between the AI’s contribution and the researcher’s judgement visible—and gives the researcher somewhere meaningful to intervene.

Architecture cannot make an interpretation true. A system can check whether an AI output has the required structure, whether it refers to permitted evidence, and whether its history has been retained. Those are important forms of technical control, but they do not establish that an interpretation is credible in relation to the participant or the research purpose. That remains an interpretive judgement.

The contribution of the architecture is therefore more modest, but still important. It makes the process through which an interpretation arose more visible, contestable and recoverable. It turns “the researcher remains responsible” from a general assurance into a set of identifiable decisions within the research process. The architecture does not, by itself, determine the status an AI output should have.

A proposal is not a finding

That distinction travels beyond the LCT.

Researchers might ask a generative AI system to summarise literature, suggest codes, identify possible themes, compare cases or develop alternative explanations. These are not merely hypothetical uses. Schroeder and colleagues, drawing on interviews with 20 qualitative researchers, found interest in using language models across stages of research alongside concerns about performance, appropriateness, ethics and participant interests.

Each of these activities may be useful. Yet the resulting text, table or classification does not arrive with its epistemic status attached. The interface presents an answer. It is the research method that must determine whether that answer is a prompt for further thought, a candidate interpretation, an item requiring verification—or something that should not enter the research record at all.

The risk is not confined to an AI producing a plainly false statement. A fluent and plausible response can move quietly from suggestion to assumption, particularly when it accords with what the researcher already expects to find. Once incorporated into notes, coding or analysis, its provisional origin may become difficult to see.

Calling an output a “proposal” helps, but the label alone is insufficient. The surrounding process must allow the researcher to inspect the evidence on which the proposal depends, understand what the AI was asked to do, alter or reject the result, and record what was ultimately accepted. Rejection is not a failure of the system; it must be one of the outcomes the process permits.

This also makes me cautious about the familiar assurance that a human remains “in the loop”. If a researcher is presented with hundreds of AI-generated classifications, without accessible evidence or sufficient time to examine them, human presence may amount to little more than administrative confirmation. Meaningful review requires both authority and the practical conditions in which to exercise it.

Lynette Pretorius and Chris Pretorius have conceptualised a researcher–participant–ChatGPT triad in which interpretation and power are negotiated. The LCT takes a more deliberately bounded route, but the comparison reinforces an important point: participation and authority need to be designed, not merely assumed.

The relevant decision should therefore be made before the AI is invited into the process: what status can its output have, and what must occur before that status changes? For one task, an output may remain a private prompt for reflection. For another, it may become a candidate code that requires comparison with source material. Any change in status requires a separate act of researcher judgement.

Treating AI output as provisional does not diminish its possible value. That provisionality provides space for AI to extend the researcher’s analytic reach while preserving the responsibility to decide what the evidence can support.

AI may contribute to an interpretation. It cannot inherit responsibility for it.

The model is not the method

The LCT has now been developed across different operating systems, using local and cloud-based AI, multiple language models, and several ways of accessing or coordinating them. At first glance, this may look like unnecessary technical complexity. Instead, it has helped clarify which parts of the work must remain stable when the technology changes.

I describe the approach as methodologically model-agnostic, but operationally model-specific.

Model-agnostic does not mean that the model is irrelevant. Every actual execution uses a particular model and version, in a particular environment, with particular instructions. Those details can affect the output and should be recorded. Different models may produce different proposals from the same evidence.

Steps in the implimentation stage, showing most stable (research design, methods) to least stable (AI model and version)
Figure 2. Methodologically model-agnostic; operationally model-specific. The research problem, method and governance should remain comparatively stable, while environments and models can change. Researcher responsibility, documentation and provenance apply across every layer.

The method, however, should not belong to any one of them. Replacing a cloud model with a locally hosted one—or moving from one interface to another—should not alter the underlying distinctions between evidence and proposal, proposal and approval, or technical and interpretive validity. If those distinctions disappear when the product changes, they were properties of the product rather than commitments of the research method.

The same issue became visible while developing this article. Work distributed across macOS and Windows 11, and involving more than one AI environment, could not safely depend on a single conversation remembering what had already been decided. We therefore separated the shared principles guiding the work, the current state of this particular project, and the instructions needed by different AI portals. The file arrangements are not the important part. The larger point is that the durable knowledge of a collaboration should reside in inspectable artefacts under the researcher’s control.

A research collaboration that depends on one conversation remembering everything has not yet become independent of the model.

This also shapes how I understand AI-assisted creation. It is reasonable to be concerned that creating through natural-language dialogue can encourage people to accept an apparently successful output without understanding or testing it. But the capacity to create something that a researcher could not have built alone does not, by itself, make the process careless.

The problem is not that researchers can now create with systems they could not have built alone; the problem arises when increased creative reach is mistaken for reduced responsibility.

In developing the LCT, responsibility is expressed through requirements that preceded the AI implementation, work divided into stages, inspection and testing of outputs, records of decisions and failures, and the ability to reconstruct or restart the work. These practices do not guarantee that every decision is correct. They make the decisions available for scrutiny and correction.

Researchers do not need to become software engineers to work this way. They do need enough understanding to identify where evidence resides, what authority the AI has, what has been retained, and whether the work could survive a change of model, portal—or collaborator.

The LCT is not yet in full production, and I do not want this account of its architecture to imply otherwise. Key methodological and architectural questions have nevertheless been resolved sufficiently for production to be a foreseeable next stage, rather than an undefined hope. This is an account of a research instrument approaching production, not evidence from its completed deployment.

Can you identify where AI sits in your research?

The LCT is one case, arising from one qualitative research method. The architecture will not transfer unchanged to every discipline or research question. The questions it exposes, however, provide a useful test for other forms of AI-enabled research.

If you are considering using generative AI in your own work, try drawing where it sits in the research process. Then ask:

  1. Where does the evidence enter?
  2. What does the AI receive?
  3. What is the generative AI system allowed to infer or generate?
  4. What is retained?
  5. What is checked, and how?
  6. Who can reject or revise the result?
  7. What becomes the authoritative version?
  8. Would the method still make sense if the model changed tomorrow?

Difficulty answering one of these questions is not necessarily a reason to abandon the proposed use of AI. It shows where further methodological or architectural work may be needed—and where an apparently simple interaction with an AI system may conceal a consequential research decision.

I began this article with the question “Which AI should I use?” Model choice does matter. It can affect capability, privacy, cost, reproducibility and the proposals a researcher receives. But it comes after a more fundamental question: what place, if any, should AI occupy within this research method?

For researchers, AI literacy is therefore not loyalty to, or fluency with, a particular system. It is the capacity to make deliberate and defensible choices about AI’s place in the research process—and to retain responsibility for the interpretations that follow.

The method should survive the model.

Origin of this article

This article grew from my HERDSA presentation, Methodological innovation in adult learner choice-making research: Locally deployed multi-agent AI systems for interview augmentation, delivered on 7 July 2026. I thank Lynette Pretorius for suggesting that I develop the presentation’s contribution into this post for The AI Literacy Lab.

AI acknowledgement

This post was developed through iterative dialogue with OpenAI’s ChatGPT and Codex. I used these systems as dialogic, analytical, drafting and editorial partners: to test the central argument; question the relationship between the research method and its technical architecture; develop and revise the structure and prose; maintain an inspectable project record across conversations and platforms; and assist with source discovery, evidence checking and citation preparation. Codex also prepared the two explanatory diagrams and drafted the figure captions and alternative text from the documented LCT architecture and our agreed visual brief. OpenAI’s built-in image-generation system produced the featured image from a specification developed through the same dialogue. I supplied the primary Learning Choice Tracker records, selected and reviewed the visual set, reviewed and revised successive drafts, and determined the method, claims, interpretations and final wording. Responsibility for the published text and visual interpretation remains mine.

Sources and further reading

Leave a comment