class: title-slide, left, bottom <span style="font-size: 50px;"> Practical AI for Science: An Orientation </span> <br> <br><br><br><br><br><br><br><br> <br><br> <table style="border: none; border-collapse: collapse; margin-left: 0; margin-top: -50px; float: left;"> <tr style="border: none; background: transparent; line-height: 0.8;"> <td style="border: none; border-right: 1px solid #6ca3d9; padding: 2px 5px; background: transparent;"><strong>Ashley I Naimi, PhD</strong></td> <td style="border: none; padding: 2px 5px; background: transparent;">Dept of Epidemiology</td> </tr> <tr style="border: none; background: transparent; line-height: 0.8;"> <td style="border: none; border-right: 1px solid #6ca3d9; padding: 2px 5px; background: transparent;">Professor</td> <td style="border: none; padding: 2px 5px; background: transparent;">Emory University</td> </tr> </table> <div style="line-height: 1.2;">
<a href="mailto:ashley.naimi@emory.edu">ashley.naimi@emory.edu</a> <br>
<a href="https://ainaimi.github.io/">https://ainaimi.github.io/</a> </div> <img src="images/qr_mice_tigers.svg" class="qr-code" alt="QR code linking to the Mice and Tigers Substack"> --- name: disclosure # AI Use Statement <br> .font160[ These slides were prepared in part using Claude-Code Fable 5, which drafted code for slide layout, key visuals, and slide style files. All content was reviewed and approved by me. The framing, outline, and arguments are my own, drawn from my course materials and decisions made across working sessions. ] --- name: outline # Outline <br> .font160[ * "Why?" and "What?" * Three ways to engage with AI * Delegation and stakes ladders * An example protocol * Some final considerations * The dangers of deskilling ] --- # Preamble <br> .font160[ This talk is about how to use AI in/for research. Not ethical, legal, social, environmental, and other issues. - bias, consent, accountability for harm - copyright and training data, liability, regulation - labor markets, misinformation, concentration of power - electricity, water, hardware - who gets access; who bears the costs - dependence, attachment, mental health ] --- name: why-evidence class: section-one # Why this talk? <table class="s1-evidence s1-agent-use" aria-label="Regular use of command-line AI coding assistants in selected disciplines, from a nonrepresentative survey in February and March 2026"> <colgroup><col style="width: 55%;"><col style="width: 45%;"></colgroup> <thead><tr><th scope="col"> </th><th scope="col">Regular coding-agent use</th></tr></thead> <tbody> <tr><td>Economics</td><td>39%</td></tr> <tr><td>Political science</td><td>25%</td></tr> <tr><td>Public health</td><td>6%</td></tr> <tr><td>Communication</td><td>6%</td></tr> <tr><td>Education</td><td>4%</td></tr> </tbody> </table> <div class="s1-source">Command-line assistants used more than once a week; nonrepresentative sample, Feb–Mar 2026.<br> Adapted from <a href="https://www.anthropic.com/research/coding-agents-social-sciences">Lyttelton, Massenkoff & Wilmers, 2026, Fig. 1</a>.</div> ??? And before we get into the main topics we'll discuss today, it would help us orient if we understood a little bit about why I put this talk together. Clearly, AI has been at the forefront of the scientific and social discourse over the past several years, and in academia many are adopting it as a tool for research, teaching, and productivity. However, in public health, the number of people engaging with AI is considerably lower than other fields, like economics. This is probably due to a number of reasons, including how the nature of the work we do differs from other fields CLICK --- name: why-local class: section-one # Why this talk? <p class="s1-kicker">What I heard at the RSPH faculty meeting:</p> <div class="s1-barrier"> <span class="s1-label">A key barrier raised in the discussion:</span> <p>Understanding what AI is and how it works.</p> </div> <br> <p class="s1-kicker">"I use AI to ... "</p> <div class="s1-tasks" role="group" aria-label="Mentioned Uses at RSPH"> <div class="s1-task">Ask questions/search</div> <div class="s1-task">Edit text</div> <div class="s1-task">Reformat material</div> </div> <p class="s1-interface">Mostly via standard web-chat interface.</p> ??? But one reason seems to be an overall lack of understanding of what AI is, how we can apply it to specific problems relevant to us as epidemiologists, biostatisticians, and behavioral, environmental, global health, or health policy and management scientists. At the full RSPH Faculty Meeting AI discussion, many of the faculty noted that they had used AI as a generic search engine to answer questions and find literature; as a tool to help refine and edit written text, and to reformat material, for example, when we have to change a manuscript to meet journal requirements. However, many were unaware of the many different ways in which we can engage with AI, and few were able to share strategies to incorporate AI prudently into their workflows. My hope is that this talk helps alleviate these knowledge and practice gaps, provide a sharper framework for understanding what AI is and how to use it, how to be careful using it, and inspire you to develop your own frameworks in a way that's most useful to you, should you choose to incorporate AI into your workflow. --- name: what-is-generative-ai class: section-one, token-orientation # What are we talking about? .font140[ Generative AI * A **large language model (LLM)** is a model fit to an enormous corpus of text (public web, books, code): it returns a prediction of the next "word-piece" (**token**). Conceptually: ] `$$\text{ChatGPT, Claude, Gemini, Copilot} \;\approx\; \hat{p}\left[\text{next token} \,\middle|\, \text{context, constraints}\right]$$` ??? To that end, it would benefit us to spend some time focusing on the high level concepts related to AI. Importantly, these concepts are not exact in the sense that they are mathematically correct. Hopefully they are useful at a conceptual level. To start, we can clarify what it is we are talking about when we say "AI". By far, the most common models that serve as the engines running an AI algorithm are large language models. A large language model is a model fit to an enormous amount of text data, such as public websites, books, code, and other repositories. These models return a prediction of the next "word-piece" or "token". Conceptually speaking, you can think of the text that comes out of the large language models underneath frontier AI company tools such as chatGPT, claude, gemini, and copilot as a token prediction, conditional on context and constraints. CLICK -- <div class="next-token-example" role="group" aria-label="Wolfram's GPT-2 example: the text so far and its five most probable continuations"> <p class="next-token-prefix">The best thing about AI is its ability to</p> <table class="next-token-table" aria-label="Reported continuation probabilities for the example text"> <colgroup><col style="width: 60%;"><col style="width: 40%;"></colgroup> <thead><tr><th scope="col">Continuation</th><th scope="col">Probability</th></tr></thead> <tbody> <tr><td>learn</td><td>4.5%</td></tr> <tr><td>predict</td><td>3.5%</td></tr> <tr><td>make</td><td>3.2%</td></tr> <tr><td>understand</td><td>3.1%</td></tr> <tr><td>do</td><td>2.9%</td></tr> </tbody> </table> </div> <div class="s1-source">GPT-2 example; only the top five continuations shown. Adapted from <a href="https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/">Wolfram, 2023</a>.</div> ??? So, for example, we might have a sequence of words such as "The best thing about AI is its ability to ..." as the context in the model, and an LLM would use this to generate predicted probabilities of the next "word piece". It would then select the word piece with the highest probability, in this case, "learn". --- name: what-are-we-talking-about class: section-one, agency-orientation # What are we talking about? <p class="s1-kicker">The problem: "LLMs kind of suck." (Hadley Wickham)</p> <br> <br> <table class="s1-evidence s1-bare-model" aria-label="Tasks a bare large language model handles unreliably, each paired with the one-line R command that performs it exactly"> <colgroup><col style="width: 45%;"><col style="width: 55%;"></colgroup> <thead><tr><th scope="col">Ask a bare LLM to…</th><th scope="col">…or run one line of R</th></tr></thead> <tbody> <tr> <td>Count the n's in "unconventional" <span class="s1-detail">100 runs: answers 1 (×14), 2 (×76), 3 (×10), not 4.</span></td> <td><code>str_count("unconventional", "n") # 4</code></td> </tr> <tr> <td>Multiply two large numbers <span class="s1-detail">20598162 x 83106206 = something reasonable, but wrong.</span></td> <td><code>20598162 * 83106206 # exact</code></td> </tr> <tr> <td>"Delete the CSV files in my directory" <span class="s1-detail">"LLMs can't do anything… They're stuck in a box."</span></td> <td><code>file.remove(list.files(pattern = ".csv"))</code></td> </tr> </tbody> </table> <div class="s1-source"><a href="https://tidydesign.substack.com/p/a-no-bullshit-guide-to-llms">Wickham, "A no-bullshit guide to LLMs," Tidy Design, Sep 5, 2026</a> (from a useR! 2025 keynote).</div> ??? The problem is that LLMs, to use a turn of phrase from Hadley Wickham, kind of suck. If you ask a bare LLM to count the number of letters in a word, such as the number of "n"'s in "unconventional" they will often get it wrong. If you ask a bare LLM to multiply two large numbers, they might give you something that looks very reasonable, but can often be wrong. You might ask an LLM to delete files in a folder on your computer, but they can't do that because they're "stuck in a box" so to speak. In contrast, each of these tasks can be easily and reliably accomplished by any number of programs. For instance, we can do each of these easily with a single short line of code in R or Python. CLICK -- <p class="s1-takeaway">LLMs, alone, can't really do these well. But they can delegate: the model calls ordinary programs and reads their results back → agentic AI.</p> ??? So LLM's alone may not reliably do these things very well. But LLMs can delegate. That is, the LLM can be used to write code that calls ordinary programs on your computer behind the scenes to do these things for you. For instance, if you open up Claude on the web and ask it to multiply 20,598,162 * 83,106,206, it will very likely use a background program such as python to compute this number for you, and not try to generate it by predicting tokens from the LLM. This ability to delegate tasks to software that is "outside of the model" so to speak, is an important part of Agentic AI. --- class: section-one, agency-orientation # What are we talking about? <div class="agency-comparison agency-single"> <div class="agency-mode"> <h2>Agentic use</h2> <p class="agency-example">“Recreate this figure in the project.”</p> <svg class="agency-flow" viewBox="0 0 535 140" role="img" aria-labelledby="agent-flow-title agent-flow-desc"> <title id="agent-flow-title">Model-directed tool-use loop</title> <desc id="agent-flow-desc">The human supplies a goal. The AI selects a step, a tool acts, and the result returns to the AI to inform its next step. Human-set limits and review points bound this loop.</desc> <rect class="agency-node agency-human" x="1" y="32" width="111" height="62" rx="3"/> <rect class="agency-node" x="163" y="32" width="185" height="62" rx="3"/> <rect class="agency-node" x="399" y="32" width="135" height="62" rx="3"/> <path class="agency-arrow" d="M 120 63 H 146"/> <polygon class="agency-arrowhead" points="144.5,57.75 155,63 144.5,68.25"/> <path class="agency-arrow" d="M 356 63 H 382"/> <polygon class="agency-arrowhead" points="380.5,57.75 391,63 380.5,68.25"/> <path class="agency-arrow" d="M 466 100 V 128 H 255 V 109"/> <polygon class="agency-arrowhead" points="249.75,110.5 255,100 260.25,110.5"/> <text x="56" y="70">Goal</text> <text x="255" y="70">AI selects step</text> <text x="466" y="70">Tool acts</text> <text class="agency-result" x="358" y="119">Result</text> </svg> </div> </div> <br> <p class="agency-distinction">Agentic versus "non-Agentic": modes of use, not types of interface. <br><br> Agentic AI = Delegation by LLM + Autonomy.</p> <div class="s1-source">Working distinction: <a href="https://www.anthropic.com/engineering/building-effective-agents">Anthropic, <em>Building effective agents</em>, 2024</a>. Schematics show simplified examples.</div> ??? By "agentic AI", I mean the ability of a model to direct a working loop like the one on this slide. Here, the user supplies a goal such as recreating a figure. The model selects a step toward that goal. A tool, such as an R script or shell command, carries out that step. The result returns to the model, which reads it and selects the next step. The loop repeats until the goal is met, it hits a limit, or it hits some review point set by the user. There are two components here that separate this from non-agentic use. The first is delegation: the model hands work to ordinary programs instead of predicting the answer token by token. The second is autonomy: within the loop, the model decides what to do next based on output. The user sets the goal and (possibly) the boundaries, but the model has say in how and whether things proceed. Importantly, agentic and non-agentic are modes of use, not interface or tool types. --- # Using Agentic AI for what? .font160[ * Statistical programming. * Manuscript development. * Internal manuscript review. * Literature surveillance. * Grant development. * Project management. * Knowledge management. ] ??? Writing, debugging, and refactoring R/Python code for simulation studies and data analyses, keeping pipelines reproducible, and making them shareable / open source via git and GitHub. Evaluating drafts of Methods and Results sections, formatting to a target journal's style, argument tightening, and reference management (e.g., changing citation styles from one journal to another). I've built a multi-agent "referee" that runs an extensive adversarial review of our drafts before submission: independent specialist reviewers, a falsification pass, verification and deduplication, then an adjudicated final report. For important findings I sometimes run the same review across two vendors' models (openAI and Anthropic) as an independence check. This has turned out to be very useful (but token heavy, so a bit pricy), and I am working on perfecting/testing it. A small agent I wrote searches PubMed, arXiv, medRxiv, and Semantic Scholar weekly, scores every new paper against a profile built from my own papers, and emails me a ranked digest. It runs unattended in the cloud via API, and costs about a dollar a week. Drafting project abstract and narratives, budget justifications based on a specific aims page and research proposal, constructing timelines and gantt charts, and running a "brutal critic" on aims pages / grant. My project folders contain machine-readable status files, and I have an agent that aggregates these into a dashboard, checks activity against my quarterly commitments, and runs my Sunday weekly planning session. AI helps me manage a linked note system (via a software tool called Obsidian, a system I've been using since well before AI). I use it to create structured notes from papers I read, and end-of-session distillation of concepts and frameworks worth keeping, and funneling into the right note. Makes past reading and past projects machine retrievable when writing the next paper or grant. It has been of tremendous use and great for generating new ideas from stuff I've read / done in the past. --- name: engage-overview class: center, middle, engagement-section # Three ways to engage with AI <div class="engage-route-labels"><span>Web interfaces</span><span>Graphical workspaces</span><span>Terminal applications</span></div> ??? So far we have talked about what AI is. Next, we address a more practical, about how we actually engage with it? There are many ways in which one can use AI, but three are general enough to mention here. The main difference between these three is how much of your computer the AI can access. --- name: engage-web class: engagement-section # 1. Web interfaces <p class="engage-subtitle">Conversation, uploaded material, and connected sources</p> <div class="engage-layout"> <figure class="engage-visual"> <img src="assets/screenshots/chatGPT-web.png" alt="ChatGPT open in a browser at chatgpt.com: a Chat/Work toggle with Chat selected, an Ask ChatGPT composer, and suggestion shortcuts. The sidebar is collapsed and no conversation is open."> <figcaption>ChatGPT web</figcaption> </figure> <dl class="engage-details"> <dt>Benefit</dt><dd>No local installation;<br>suited to conversational iteration.</dd> <dt>Tradeoff</dt><dd>Limited source file access<br>and output return.</dd> </dl> </div> <div class="engage-source">Capabilities: <a href="https://learn.chatgpt.com/docs/web">OpenAI, ChatGPT on the web</a>; <a href="https://learn.chatgpt.com/docs/projects">Projects and chats</a>. Checked Sep 13, 2026. Uses are suggested examples.</div> ??? The first is the web interface, the one most of us already use. There's nothing to install, and it's well suited to conversational back and forth with uploaded documents and connected sources. The tradeoff is reach, in that it has limited access to the files on your machine, and moving its output back into your project has to be done manually. --- name: engage-workspace class: engagement-section # 2. Graphical workspaces <p class="engage-subtitle">Claude Cowork: delegate work across files and tools</p> <div class="engage-layout"> <figure class="engage-visual"> <img src="assets/screenshots/engage-cowork.png" alt="Claude Cowork desktop session showing a greeting, a model-usage panel, and a composer with a connected local project folder."> <figcaption>Claude Cowork</figcaption> </figure> <dl class="engage-details"> <dt>Benefit</dt><dd>Multi-file delegation<br>without a terminal.</dd> <dt>Tradeoff</dt><dd>Restricted scope access; operates only in assigned folder.</dd> </dl> </div> <div class="engage-source"><a href="https://support.claude.com/en/articles/13345190-get-started-with-claude-cowork">Anthropic, Cowork guide</a>; <a href="https://learn.chatgpt.com/docs/get-started-with-work">OpenAI, ChatGPT Work</a>. Checked Sep 13, 2026. Uses are suggested examples.</div> ??? The second is a graphical workspace, such as Claude Cowork. You assign it a project folder, and it can read and edit the files in that folder, with no terminal or shell interface required. The tradeoff is scope: its access to your local filesystem is generally limited to the folders you explicitly give it access to. --- name: engage-terminal class: engagement-section # 3. Terminal applications <p class="engage-subtitle">Claude Code</p> <div class="engage-layout"> <figure class="engage-visual"> <img src="assets/screenshots/engage-terminal-claude.png" alt="Claude Code running in a terminal: startup screen with version, model, and working directory, with auto mode enabled. No task has been run."> <figcaption>Claude Code</figcaption> </figure> <dl class="engage-details"> <dt>Benefit</dt><dd>Runs scripts, git,<br>tests, and batch jobs.</dd> <dt>Tradeoff</dt><dd>Command literacy;<br>control write access<br>and execution.</dd> </dl> </div> <div class="engage-source"><a href="https://code.claude.com/docs/en/overview">Anthropic, Claude Code overview</a>; <a href="https://learn.chatgpt.com/docs/codex/cli">OpenAI, Codex CLI</a>. Checked Sep 13, 2026. Uses are suggested examples.</div> ??? The third is a terminal-based agent, such as Claude Code running from the command line. This gives it potentially much broader reach: it can run scripts, use Git, execute tests, launch batch jobs, and invoke other command-line tools on your machine. The tradeoff is that it requires some command-line literacy, and you need to pay attention to what the agent is allowed to write and execute. --- name: ladders-overview class: center, middle, ladders # Delegation and stakes ladders <p class="ladders-divider-note"> <em>Will delegating pay?</em> · <em>What price if it's wrong?</em></p> ??? Once we determine how we will engage with AI, a more difficult question involves judgement. I find it use ful to separate these judgements into at least two inter-related categories. First: is delegating a task or set of tasks to AI actually worth the effort? Second: what price do we pay if the delegated work is wrong? What's at stake? --- name: delegation-question class: section-one, ladders # Is it worth delegating? <p class="s1-kicker">Before delegating a task to AI, evaluate whether:</p> `$$\text{specify} \;+\; \text{supervise} \;+\; \text{verify} \;+\; \text{repair} \;\;<\;\; \text{do it yourself}$$` <div class="s1-source">Working guidance: <a href="https://www.anthropic.com/engineering/building-effective-agents">Anthropic, <em>Building effective agents</em>, 2024</a> ("find the simplest solution possible"). A heuristic, not a measured model.</div> ??? Of course, one can use AI for almost anything we might do sitting at our computers, but whether it's worth our time and energy to do so should be a question considered explicitly before embarking on carrying out a project with AI. Generally, we should always consider whether the work of specifying and supervising an AI, verifying it's output, and repairing what might have been done incorrectly is overall less than the work that would result from us simply doing it ourselves. CLICK -- <br> .font130[ When a task recurs (amortization) and/or when verification is simple, delegation usually pays off. ] ??? One helpful heuristic lies in considering amortization over time, and our ability to verify the work that AI does. We'll see a specific example of this in a couple of slides, but generally when a task is recurring and easy to verify, it may be a good candidate for AI to work on with you. --- name: delegation-ladder class: section-one, ladders # The delegation ladder <svg class="ladder-stair" viewBox="0 0 1120 400" role="img" aria-labelledby="lad-stair-title lad-stair-desc"> <title id="lad-stair-title">The delegation ladder</title> <desc id="lad-stair-desc">Four rungs of increasing delegation: assistant, collaborator, delegate with structure, delegate with autonomy. Beneath the rungs, the human role shifts from operating to approving to observing.</desc> <rect class="lad-node" x="10" y="210" width="260" height="90"/> <rect class="lad-node" x="282" y="150" width="260" height="150"/> <rect class="lad-node" x="554" y="90" width="260" height="210"/> <rect class="lad-node" x="826" y="30" width="260" height="270"/> <text class="lad-name" x="140" y="246">1. Assistant</text> <text class="lad-ex" x="140" y="278">"explain this estimator"</text> <text class="lad-name" x="412" y="186">2. Collaborator</text> <text class="lad-ex" x="412" y="218">"critique this draft,</text> <text class="lad-ex" x="412" y="242">with my documents"</text> <text class="lad-name" x="684" y="122">3. Delegate,</text> <text class="lad-name" x="684" y="150">with structure</text> <text class="lad-ex" x="684" y="182">"run the analysis under</text> <text class="lad-ex" x="684" y="206">a plan I reviewed"</text> <text class="lad-name" x="956" y="62">4. Delegate,</text> <text class="lad-name" x="956" y="90">with autonomy</text> <text class="lad-ex" x="956" y="122">"standing weekly</text> <text class="lad-ex" x="956" y="146">literature digest"</text> <text class="lad-axis lad-left" x="10" y="326">your role</text> <path class="lad-line" d="M 10 340 H 1078"/> <polygon class="lad-arrowhead" points="1076,334.75 1086.5,340 1076,345.25"/> <text class="lad-role" x="140" y="374">you operate</text> <text class="lad-role" x="412" y="374">you operate</text> <text class="lad-role" x="684" y="374">you approve</text> <text class="lad-role" x="956" y="374">you observe</text> </svg> ??? The question of whether it's worth our time to assign a project, or components of a project, to AI leads to the concept of a delegation ladder, capturing questions about roughly WHAT and HOW MUCH we delegate to AI? On the leftmost side of this ladder or staircase, we delegate the least, using AI as little more than a search engine or a tool to compile and organize information. At the other end of the spectrum, we give AI the autonomy to make decisions without our input or approval. -- <div class="ladder-dims"> <p><span class="lad-q">Who chooses the next step?</span> you at every turn → you at checkpoints → the model, with structure and checks → the model</p> <p><span class="lad-q">The unit of work handed over?</span> a question → a working session → a bounded task → a standing process</p> <p><span class="lad-q">When verification happens?</span> every response → at checkpoints → after the fact</p> </div> <div class="s1-source">(Graded autonomy as a design choice): <a href="https://apps.dtic.mil/sti/citations/ADA057655">Sheridan & Verplank, 1978</a>; <a href="https://www.sae.org/standards/content/j3016_202104/">SAE J3016</a>; <a href="https://arxiv.org/abs/2311.02462">Morris et al., 2024</a>.</div> ??? This ladder can be conceptualized in many ways. For instance, we might ask "who chooses the next step in the work process?", and increasing amounts of delegation might lead to a progression such as "you, every time", "you at prespecified checkpoints", the model, with structure and checks, and the model with structure, but no checks. As an example of full delegation with structure, consider a tool that I recently built for myself, which is a literature search agent. --- name: autonomy-example class: section-one, ladders # Delegate with Autonomy: An Example <img class="lad-strip" src="assets/screenshots/lit-search-header.png" alt="Header of the Weekly Literature Digest email, September 6, 2026: 457 papers searched, 66 passed threshold (at least 6 of 10), top 20 shown; full scored list of 66 papers attached as an HTML file."> <svg class="lit-flow" viewBox="0 0 1120 200" role="img" aria-labelledby="lit-flow-title lit-flow-desc"> <title id="lit-flow-title">The literature agent's pipeline</title> <desc id="lit-flow-desc">Four sources — PubMed, arXiv, medRxiv, Semantic Scholar — feed deduplication and a keyword pre-filter; a language model scores each surviving paper against the presenter's publication profile; the top twenty arrive as a Sunday email. Counts from the September 6, 2026 run: 457 searched, 66 passed the threshold, 20 shown.</desc> <rect class="lad-node" x="6" y="26" width="300" height="78"/> <text x="156" y="58">PubMed · arXiv</text> <text x="156" y="84">medRxiv · Semantic Scholar</text> <path class="lad-line" d="M 306 65 H 332"/> <polygon class="lad-arrowhead" points="330,59.75 340.5,65 330,70.25"/> <rect class="lad-node" x="344" y="26" width="222" height="78"/> <text x="455" y="58">deduplicate +</text> <text x="455" y="84">keyword pre-filter</text> <path class="lad-line" d="M 566 65 H 592"/> <polygon class="lad-arrowhead" points="590,59.75 600.5,65 590,70.25"/> <rect class="lad-node" x="604" y="26" width="290" height="78"/> <text x="749" y="58">LLM scores each paper</text> <text x="749" y="84">vs. my user profile</text> <path class="lad-line" d="M 894 65 H 920"/> <polygon class="lad-arrowhead" points="918,59.75 928.5,65 918,70.25"/> <rect class="lad-node" x="932" y="26" width="182" height="78"/> <text x="1023" y="58">top-20 digest,</text> <text x="1023" y="84">emailed Sunday</text> <text class="lit-count" x="156" y="138">457 papers</text> <text class="lit-count" x="749" y="138">66 pass ≥ 6/10</text> <text class="lit-count" x="1023" y="138">20 shown</text> <text class="lad-axis" x="560" y="182">runs unattended in the cloud, weekly · $0.50–0.90 per run · remembers what it has already shown</text> </svg> ??? One problem that I often encounter given my eclectic research interests across a wide range of fields is keeping up with a literature in general and on specific topics. At some point, this led to a mess of journal TOCs, mailing lists from google scholar, pubmed, the arxiv, and medrxiv that was simply out of hand. So i created an online literature search agent that runs weekly in the cloud, searchers all the databases that I would be searching typically, and that uses an LLM structure with information that I give it in the form of papers that I like, papers that I've written, authors that I'd like to follow, and other information, and ranks a set of papers published in the preceding week using a relevance scale that is derived from the information I provided the LLM. Every sunday, I get an email with the top 20 ranked papers on topics that I am interested in, with a full scored list attached as an html file. This literature search agent decides completely on its own (based on the context I provided) which papers are most relevant, and which less so, and then notifies me by email of what it found. --- name: stakes-ladder class: section-one, ladders # The stakes ladder <svg class="stakes-stair" viewBox="0 0 1120 436" role="img" aria-labelledby="stakes-stair-title stakes-stair-desc"> <title id="stakes-stair-title">The stakes ladder</title> <desc id="stakes-stair-desc">Four tiers of rising stakes: routine (private and reversible), consequential (informs your thinking and drafts), high (enters the record), critical (irreversible, or protected). Beneath the steps, the matching rule for each tier, from delegate freely to mostly don't.</desc> <rect class="lad-node" x="10" y="210" width="260" height="90"/> <rect class="lad-node" x="282" y="150" width="260" height="150"/> <rect class="lad-node" x="554" y="90" width="260" height="210"/> <rect class="lad-node lad-human" x="826" y="30" width="260" height="270"/> <text class="lad-name" x="140" y="246">1. Routine</text> <text class="lad-ex" x="140" y="278">private and reversible</text> <text class="lad-name" x="412" y="186">2. Consequential</text> <text class="lad-ex" x="412" y="218">informs your thinking</text> <text class="lad-ex" x="412" y="242">and drafts</text> <text class="lad-name" x="684" y="126">3. High</text> <text class="lad-ex" x="684" y="158">enters the record,</text> <text class="lad-ex" x="684" y="182">affects your reputation</text> <text class="lad-ex" x="684" y="182">as a scholar</text> <text class="lad-name" x="956" y="62">4. Critical</text> <text class="lad-ex" x="956" y="94">irreversible,</text> <text class="lad-ex" x="956" y="118">or protected</text> <text class="lad-axis lad-left" x="10" y="326">the rule</text> <path class="lad-line" d="M 10 340 H 1078"/> <polygon class="lad-arrowhead" points="1076,334.75 1086.5,340 1076,345.25"/> <text class="lad-role" x="140" y="374">delegate freely</text> <text class="lad-role" x="412" y="374">ground in sources;</text> <text class="lad-role" x="412" y="398">spot-check; log</text> <text class="lad-role" x="684" y="374">plan review;</text> <text class="lad-role" x="684" y="398">independent verification;</text> <text class="lad-role" x="684" y="422">provenance; disclose</text> <text class="lad-role" x="956" y="374">mostly: don't</text> </svg> <div class="ladder-dims dims-tight"> <p><span class="lad-q">What breaks if it's wrong?</span> information → a program you run once → your computer → the internet</p> <p><span class="lad-q">Is it verifiable?</span> easily → with some effort → with great effort → unverifiable</p> <p><span class="lad-q">Is it reversible?</span> completely → cheaply → with great effort → completely irreversible damage</p> </div> <div class="s1-source">"Chatbots ... cannot be responsible for the accuracy, integrity, and originality of the work." — <a href="https://www.icmje.org/recommendations/browse/roles-and-responsibilities/defining-the-role-of-authors-and-contributors.html">ICMJE Recommendations</a>. Same tiered shape as the <a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai">EU AI Act</a> and <a href="https://www.imdrf.org/documents/software-medical-device-possible-framework-risk-categorization-and-corresponding-considerations">FDA/IMDRF</a> risk categories.</div> ??? The second ladder focuses on a different question: what does it cost if the work is wrong? What's at stake? The stakes ladder starts with work that is private and reversible; through consequential work that informs my thinking and my drafts; to high-stakes work that enters the scientific record and becomes a part of my reputation; and finally a critical tier level, where an error is is high stakes and irreversible. The three questions at the bottom locate a task on the ladder: what breaks if it is wrong, how hard the work is to verify, and whether a mistake can be undone. Each tier then carries its own rule, along the bottom of the staircase. Routine work: delegate freely. Consequential work: ground it in sources, spot-check it, keep a log. High stakes: review the plan before it runs, verify independently, keep provenance, and disclose the AI's role. Critical: mostly, don't. And here is how the two ladders work together, which is the one sentence I would like you to keep: the rung you choose on the delegation ladder has to be justified by the tier the task occupies on the stakes ladder. High delegation on low-stakes work is where the payoff lives. High delegation on high-stakes work is where the trouble lives. As the stakes rise, we climb down the delegation ladder, even when the model could probably do the task. The ICMJE line at the bottom anchors the top tier: a chatbot cannot be responsible for the accuracy, integrity, and originality of the work. Accountability cannot be delegated. That is why, at the top of this ladder, the autonomy stays with us. --- name: protocol-overview class: center, middle, ladders # A protocol and example <p class="ladders-divider-note">One small project, with folders, spec, prompts, checks, and a handoff.</p> ??? --- name: proto-problem class: section-one, ladders # Example Application <br> <br> .font140[ * Quick exploration of new methods for effect modification (conditional average treatment effect) estimators. * The estimators, called **DR-** and **R-learner**, using the same super learner for outcome/exposure. * They also BOTH use LOESS for the actual learners. * The true CATE function is known by construction: τ(x) = 2·sin(x), best visualized (later). * The question: if I use exactly the same modeling strategies for both the **DR-** and **R-learner**, will I see a difference? ] <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: proto-house class: section-one, ladders # The Folder Setup Files and folders that begin with "underscore" are AI specific: <pre class="proto-tree"> ai-protocol-example/ ├── _corpus/ sources material for the project ├── _notes/ the project's memory: notes and closeouts ├── _specs/ instructions for AI to carry out the work <span class="tree-red">├── code/ R code for simulation and analysis</span> <span class="tree-red">├── output/ output that the code produced</span> <span class="tree-red">├── figures/ figures generated from the code</span> ├── PROMPTS.md example AI prompts <span class="tree-red">└── README.md outline of project folder</span> </pre> <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: proto-spec class: section-one, ladders # Write the spec before any prompt <p class="s1-kicker">The spec "specifies" the shape of the project, what it should look like, and thus what the AI should do to move the project forward (and how). Excerpt:</p> <pre class="proto-tree"> ## Question How do the DR-learner and the R-learner compare when estimating a known conditional average treatment effect (CATE) function, with the nuisance modeling and smoothing held identical between them? ## Data-generating process For each of `n` independent individuals: - Covariate: `X ~ Uniform(0, 2*pi)` - Exposure: `A | X ~ Bernoulli( plogis(0.5 * sin(X)) )` - Potential outcomes: `Y0 = 5 + 0.5*X + Normal(0, 0.5)`; `Y1 = Y0 + tau(X)` - Observed outcome: `Y = A*Y1 + (1-A)*Y0` </pre> <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: proto-prompt-plan class: section-one, ladders # Prompt 1: Planning <div class="prompt-quote">Read _specs/2026-09-16-cate-comparison-spec.md, README.md, and the papers listed in _corpus/README.md. Propose a step-by-step plan to implement the spec exactly: file layout, functions, how the matched nuisance and smoothing constraints will be enforced, the assertions, and the figure. Flag anything in the spec that is ambiguous or underdetermined. Do not write any code yet.</div> <br> <br> <div class="ladder-dims"> <p><span class="lad-q">The mode is plan only: the AI tool reads and proposes a plan.</p> <p><span class="lad-q">The plan is grounded in the spec and the corpus papers.</p> <p><span class="lad-q">The human task is to review and approve the plan, or fix the spec first.</p> </div> <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: proto-prompt-execute class: section-one, ladders # Prompt 2: Execution <div class="prompt-quote">Execute the approved plan. Write the complete single-replicate pipeline as code/01_simulation.R, following the spec's DGP, estimators, metrics, and verification assertions; resolve any question about the DR-learner or R-learner against the papers in _corpus/. Ask me if anything is unclear. Save the ISE table to output/ise.csv. Then generate the figure comparing the different estimators.</div> <br> <br> <div class="ladder-dims"> <p><span class="lad-q">For bigger projects, this is be done in steps, with verification.</p> <p><span class="lad-q">The spec's parameters and assertions instruct the AI; implementation details extracted from the corpus.</p> <p><span class="lad-q">You read the code at the checkpoints, then run it yourself, correct and proceed.</p> </div> <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: proto-verify class: section-one, ladders # Verify <img class="proto-fig" src="images/cate_comparison.png" alt="Estimated CATE curves for the DR-learner (blue) and R-learner (red) overlaid on the true curve 2 times sine of x (black), with a histogram of X along the bottom. Both estimates track the truth through the interior; the R-learner flattens at both boundaries."> <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: proto-closeout class: section-one, ladders # Closeout and Resumption <div class="ladder-dims"> <p><span class="lad-q">End of session 1:</span> a dated note to _notes/: what was done and decided, what remains, and what's next.</p> <p><span class="lad-q"Session close:</span> a "closeout" function, which is a skill that saves a summary to _notes/.</p> </div> <div class="prompt-quote">Read the most recent note in _notes/. Let's continue from where it left off.</div> <div class="s1-source">Full example: <a href="https://github.com/ainaimi/ai-protocol-example">github.com/ainaimi/ai-protocol-example</a>.</div> --- name: considerations-overview class: center, middle, ladders # Some final considerations --- name: consider-data class: section-one, ladders # This example used no protected data <p class="s1-kicker">The example ran entirely on simulated data. No PHI, and no human subjects data.</p> <br> -- <div class="ladder-dims"> <p><span class="lad-q">What the example shared:</span> every prompt, spec, and data file was sent to a commercial model server. With simulated data, this is less consequential.</p> <p><span class="lad-q">What changes with PHI:</span> the same step becomes a disclosure of protected data to a third party. HIPAA and Emory policy govern that step, and it is usually prohibited.</p> <p><span class="lad-q">Before opening AI in a folder with PHI:</span> know how your data are classified, and which tools are approved for that classification.</p> </div> <div class="s1-source">Emory Office of Responsible AI, <a href="https://responsibleai.emory.edu/guidelines/understanding-your-data-security-responsibilities.html">Understanding your data security responsibilities</a>. Checked Sep 17, 2026.</div> --- name: consider-emory class: section-one, ladders # Emory policy on AI and sensitive data <div class="policy-shots"> <figure> <img src="assets/screenshots/emory-data-security.png" alt="Top of Emory's Responsible AI site: a page titled AI Data Security Responsibilities, defining sensitive data as any information that can identify individuals or is considered confidential, restricted, or proprietary."> <figcaption>responsibleai.emory.edu</figcaption> </figure> <figure> <img src="assets/screenshots/emory-copilot-launch.png" alt="Emory News story dated April 22, 2024: Microsoft Copilot AI chat service for Emory community launches."> <figcaption>news.emory.edu</figcaption> </figure> </div> <div class="ladder-dims dims-tight"> <p>Restricted data (PHI, PII, FERPA records) may only be used with Emory-approved secure AI technology. Public chatbots can retain inputs and train on them.</p> <p>Emory's Microsoft Copilot is the protected option for everyday Emory data, but not for PHI or IIHI.</p> </div> <div class="s1-source"><a href="https://responsibleai.emory.edu/guidelines/understanding-your-data-security-responsibilities.html">Emory Responsible AI, data security responsibilities</a>; <a href="https://news.emory.edu/stories/2024/04/microsoft-copilot-ai-chat-service-emory-community-launches">Emory News, Copilot launch</a>. Checked Sep 17, 2026.</div> --- name: consider-integrity class: section-one, ladders # Scientific integrity <div class="ladder-dims"> <p><span class="lad-q">Accountability:</span> AI cannot be accountable, so it cannot be an author. Only human authors can be responsible for the accuracy, integrity, and originality of the work.</p> <p><span class="lad-q">Disclosure:</span> disclosure follows venue policy. Many journals require it, NIH prohibits AI in peer review entirely. However, a defining characteristic in science is openness, and we should all be explicit about how we use AI.</p> </div> <br> <div class="s1-takeaway">The standard has always been: You must explain and defend the work without outsourcing to an AI tool.</div> <div class="s1-source"><a href="https://academic.oup.com/aje/advance-article-abstract/doi/10.1093/aje/kwag224/8789968?redirectedFrom=fulltext">Code Sharing at AJE</a>; <a href="https://www.icmje.org/recommendations/browse/roles-and-responsibilities/defining-the-role-of-authors-and-contributors.html">ICMJE Recommendations</a>; <a href="https://publicationethics.org/cope-position-statements/ai-author">COPE position on AI and authorship</a>; <a href="https://www.jmir.org/2024/1/e53164">Chelli et al., JMIR 2024</a>; <a href="https://grants.nih.gov/grants/guide/notice-files/NOT-OD-23-149.html">NIH NOT-OD-23-149</a>.</div> --- name: deskilling class: section-one, ladders # Deskilling <div class="deskill-layout"> <figure> <img src="assets/screenshots/nyt-cognitive-surrender.png" alt="New York Times headline: An M.I.T. Report Warns A.I. Is Causing Cognitive Surrender. Universities Are in a Bind. Subhead: As A.I. upends education, university leaders have been all over the map about how to respond. It can be very confusing for students."> <figcaption>The New York Times, Sep 15, 2026</figcaption> </figure> <div class="ladder-dims deskill-points"> <p><span class="lad-q">Clinicians:</span> after months of AI-assisted colonoscopy, experienced endoscopists detected fewer adenomas in unassisted procedures: 28.4% before AI exposure, 22.4% after.</p> <p><span class="lad-q">Students:</span> MIT's committee on AI in teaching describes "cognitive surrender": turning to the chatbot at the first sign of difficulty.</p> </div> </div> -- <br> <div class="ladder-dims dims-tight"> <p><span class="lad-q">A proactive stance:</span> students and faculty should be able to specify, supervise, and verify. Those abilities come from doing the work unassisted, from deep learning.</p> <p>"The one who does the work, does the learning."</p> </div> <div class="s1-source"><a href="https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00289-4/abstract">Budzyń et al., Lancet Gastroenterol Hepatol 2025</a>; <a href="https://www.nytimes.com/2026/09/15/us/universities-ai-warnings-enthusiasm.html">The New York Times, Sep 15, 2026</a>.</div> --- class: title-slide, left, bottom <span style="font-size: 50px;"> Practical AI for Science: An Orientation </span> <br> <br><br><br><br><br><br><br><br> <br><br> <table style="border: none; border-collapse: collapse; margin-left: 0; margin-top: -50px; float: left;"> <tr style="border: none; background: transparent; line-height: 0.8;"> <td style="border: none; border-right: 1px solid #6ca3d9; padding: 2px 5px; background: transparent;"><strong>Ashley I Naimi, PhD</strong></td> <td style="border: none; padding: 2px 5px; background: transparent;">Dept of Epidemiology</td> </tr> <tr style="border: none; background: transparent; line-height: 0.8;"> <td style="border: none; border-right: 1px solid #6ca3d9; padding: 2px 5px; background: transparent;">Professor</td> <td style="border: none; padding: 2px 5px; background: transparent;">Emory University</td> </tr> </table> <div style="line-height: 1.2;">
<a href="mailto:ashley.naimi@emory.edu">ashley.naimi@emory.edu</a> <br>
<a href="https://ainaimi.github.io/">https://ainaimi.github.io/</a> </div> <img src="images/qr_mice_tigers.svg" class="qr-code" alt="QR code linking to the Mice and Tigers Substack">