| Academic Programme | UG B.A. (Hons) / B.Sc. (Hons) |
| Academic Year / Term(s) | 2026-27 / Term 4 |
| Course Credits | 4 |
| Course Type | - |
| Course Category | Mandatory |
| Host Discipline | - |
| Cross-listed Discipline(s) | - |
| Pre-requisites | CS 101 Principles of Computer Science and Software Development; Mathematics and Logic (Part 2, Introduction to Statistics) |
| Instructor(s) | Navin Kabra |
| Instructor Email Id(s) | navin@smriti.com |
| Faculty Associate (FA) | Devdatta Kathale; Ashvina Kulkarni |
| FA Email Id | Devdatta.Kathale@nayanta.edu.in; Ashvina.Kulkarni@nayanta.edu.in |
This course teaches students to think carefully about data and AI: to ask the right questions, find or collect the right evidence, choose an appropriate analysis, and report what they find honestly.
The focus is on three questions:
AI appears in three ways. As a subject of study: how do these systems learn from data, and where do they fail? As a programming tool: students use an LLM to write their code throughout the course. And as an analysis tool: to be able to quickly do many different analyses, to quickly see the consequences of different choices, and in general do a broader exploration of the design space.
Students will have a ChatGPT Plus subscription for the duration of the course, and will use it in almost every session. No prior experience with pandas, matplotlib or machine learning is assumed.
The course aims to give students:
By the end of this course, students will be able to:
The course runs over six weeks, in eighteen sessions: twelve of 90 minutes and six of two hours. A seventh week is used for assessment.
Each week has two 90-minute Tuesday sessions and one two-hour Thursday session. The Tuesday sessions introduce material, and the second one is always hands-on rather than a second lecture. Many of the Thursday sessions are labs or studios.
Specific choices worth noting:
| S.No. | Topics | Suggested Readings |
|---|---|---|
| 1 | The Data Mindset: questions, evidence and reasoning | Calling Bullshit, ch. 1–2; free course videos |
| 2 | Exploring data with AI-assisted code | — |
| 3 | Lab: environment, and one dataset end to end | — |
| 4 | When summary statistics mislead | The Art of Statistics, ch. 1–2 |
| 5 | Visualisation as argument | How Charts Lie, ch. 1–3; Data Feminism, ch. 3; Cairo’s short guide |
| 6 | Studio: duelling charts | Calling Bullshit, ch. 7; free visualisation notes |
| 7 | Where data comes from | Calling Bullshit, ch. 4; Data Feminism, ch. 6 |
| 8 | Correlation, causation and study design | The Art of Statistics, ch. 4; free causality videos |
| 9 | Lab: data you collect yourself, including data that isn’t a table | — |
| 10 | Hypothesis testing, and the replication crisis | The Art of Statistics, ch. 10–11; Ioannidis (2005) |
| 11 | What is a model? From regression to AI | The Art of Statistics, ch. 5; ISLP ch. 2 slides |
| 12 | Lab: the p-hacking sandbox | Gelman & Loken, “The Garden of Forking Paths”; FiveThirtyEight’s article and interactive |
| 13 | Evaluation: generalisation, metrics and decisions | The Art of Statistics, ch. 6; Google’s notes on overfitting and classification metrics |
| 14 | Trees, forests and clustering | ISLP ch. 2 (concepts only); official ch. 2 slides; Google’s clustering overview |
| 15 | Trustworthy AI: leakage, shift and failure modes; leakage hunt | AI Snake Oil, ch. 2–3; Princeton excerpt; Kapoor & Narayanan on leakage |
| 16 | Inside the AI: how large language models work | Wolfram, “What Is ChatGPT Doing?”; Alammar, “The Illustrated Transformer” |
| 17 | AI as your data analysis partner | — |
| 18 | Ethics, consequences and communicating data | Weapons of Math Destruction, ch. 1, 5, 6; author interview |
| Assessment day: final examination and project defence | — | |
| Free slot: Q&A and debrief | — |
Nothing on this list is required. The list is for students who want to follow something up, or who learn better from books and articles.
Two books cover much of the course:
These books/articles will be useful for particular parts of the course:
References for looking things up rather than reading:
Shorter pieces, all free, are listed against the individual sessions below.
There is no Python textbook on this list, and that is intentional. Students have ChatGPT Plus, and the course teaches them to direct it and check its output rather than to memorise syntax.
| S.No. | Component | Weightage |
|---|---|---|
| 1 | Quizzes: five short in-class MCQ quizzes, best four counted | 20% |
| 2 | Project, in pairs: proposal, analysis and written report along with all data and code | 30% |
| 3 | Project defence: individual, written, 60 minutes | 25% |
| 4 | Final examination: written, 2 hours | 25% |
Quizzes. Five quizzes of about ten minutes, at the end of a session, from Week 2 onwards. Each tests whether a concept from recent classes were learnt properly. The best four of five count, so one missed quiz costs nothing.
Project. To be done in pairs. Projects will be self-selected with the help and approval of the instructor and the faculty associates. All projects should involve an important question, some data curation and cleanup, followed by analysis, and a written report, along with all the data and the programs. The report must explain how and where AI was used in the project, and how the outputs of the AI were verified.
Project defence. After the project report is submitted and assessed by the FAs and instructor, each student receives a printed sheet of five questions written specifically about their own project, and answers them by hand in an invigilated room. Students may have their own report on the desk and nothing else. The two members of a pair get different questions and answer individually and independently.
Final examination. A two hours written, invigilated exam. Questions will typically contain some data artifact — a chart, a regression output, a description of how a sample was collected, an extract of AI-generated analysis — and will ask the student to judge it. This test will not require calculations.
There are no marks for attendance or classroom participation.
However, keep in mind that the 5 quizzes happen during class, and will not be repeated. Students who miss two or more sessions with quizzes could lose marks.
In addition, university attendance requirements apply as normal.
This course requires the use of AI. Use of AI in all aspects of the course is encouraged. Every student has a ChatGPT Plus subscription paid for by the university, and most of the code written during the course will be produced by an LLM rather than typed by hand.
However, to ensure that students actually learn things while using AI, a major part of the evaluation of the course (quizzes, project defense, and final exam) will be offline, without the use of AI or any devices.
Verification. Students are responsible for everything they submit, including the parts they did not write. “ChatGPT produced that” is not a defence for a wrong number, a misleading chart, or an analysis that does not answer the question.
Passing off another student’s work, or published text, as one’s own is dealt with under the university’s normal plagiarism policy. This is unaffected by the AI provisions above.
Sessions are grouped by teaching week. Module labels are given in brackets. All readings listed are suggested; none is required.
What data analytics is, how it is different from statistics, and why it is worth studying. The cycle: question, data, analysis, interpretation, communication. Types of data — structured, unstructured, qualitative, quantitative. How to take a vague question and make it precise enough to answer with data. Confounding, bias, and the limits of what can be learned from watching rather than intervening.
Activity: students will be given some recent newspaper headlines and work out what data and analysis would support it, and what would not.
Suggested reading: Calling Bullshit, ch. 1–2. The course videos cover the same introductory material for free.
Introduction to pandas and matplotlib, with ChatGPT writing the code. Loading, inspecting and summarising a dataset. Students write the prompt, then read the code that comes back and check whether it does what they asked for. This session establishes the working method for the rest of the course.
Getting every student’s Python and ChatGPT Plus working. Then one small dataset all the way through: load it, look at it, summarise it, produce one chart, write one sentence about what it shows. None of it is difficult; the point is that all thirty students have been through the complete loop once.
An ungraded diagnostic quiz at the end, to establish what the class remembers from CS 101 and from the statistics half-course.
Connecting the statistics the class already knows to real, messy data. Why the mean is often the wrong number to report: skew, outliers, Simpson’s paradox. Choosing a summary that answers the question being asked.
Suggested reading: The Art of Statistics, ch. 1–2.
The same data can be charted to tell very different stories. Tufte’s principles of chart design. The usual ways charts mislead: truncated axes, cherry-picked date ranges, aggregation that hides the interesting variation. What ChatGPT produces by default when asked for “a chart”, and what that default assumes on the analyst’s behalf.
Suggested readings: How Charts Lie, ch. 1–3; Data Feminism, ch. 3, “On Rational, Scientific, Objective Viewpoints”; Wilke, Fundamentals of Data Visualization, ch. 5. Cairo covers much of the same ground in this shorter illustrated article.
One dataset, class split in two. Each half is assigned a conclusion and has to build the best honest chart it can in support of it. Both sides present, and the class then argues about where the line falls between framing something well and misleading people.
Quiz 1: summaries and visualisation.
Suggested reading: Calling Bullshit, ch. 7. The free visualisation notes cover misleading axes and proportional ink.
Sampling, surveys, sensors, scraped web pages, administrative records. Selection bias, survivorship bias, measurement error. The difference between data collected deliberately and data that was found.
Case study: biased training data in deployed AI systems, for example in predictive policing and in medical imaging.
Suggested readings: Calling Bullshit, ch. 4; Data Feminism, ch. 6, “The Numbers Don’t Speak for Themselves” (complete chapter online).
Observational against experimental data. Why correlation is not causation, with enough examples that it sticks. Randomised controlled trials, natural experiments, A/B testing. Then the reverse question: given a causal question, what data would be needed to answer it? And why a model trained on observational data cannot answer it, however accurate the model looks.
Suggested reading: The Art of Statistics, ch. 4; the Calling Bullshit causality videos cover the same ideas for free.
Designing a survey or collection instrument, and the many ways it goes wrong. Then the harder case: images, audio, video, sensor readings. What is involved in turning any of those into something that can be computed with, and what gets thrown away in the process.
The project brief is issued in this session and pairs are formed.
Quiz 2: provenance and causation.
p-values and confidence intervals applied to real data. What a p-value does and does not permit a student to claim. Effect size against statistical significance. Then the replication crisis in psychology and medicine: multiple testing, the garden of forking paths, publication bias, and how all of it happens without anyone setting out to cheat.
Suggested readings: The Art of Statistics, ch. 10–11; Ioannidis, “Why Most Published Research Findings Are False”, PLoS Medicine (2005, open access).
A model is a simplified description of the world, learned from data. Linear regression as the simplest example of something that learns: fitting a line, interpreting the coefficients, looking at the residuals. Why a baseline is needed before any modelling result means anything.
Suggested reading: The Art of Statistics, ch. 5; ISLP’s chapter 2 slides are a free introduction to models, features, targets, prediction and inference.
Students are given a dataset of pure noise and asked to find a publishable result in it. Most will manage it. They then write down an analysis plan for a real dataset before looking at it, and hold themselves to it.
Project proposals are due in this session, with the analysis plan written down in advance, and each receives feedback in class.
Quiz 3: inference.
Suggested readings: Gelman and Loken, “The Garden of Forking Paths”; FiveThirtyEight, “Science Isn’t Broken”, including its p-hacking tool.
Train/test splits and cross-validation, at the level of intuition rather than formulae. Overfitting and underfitting. Choosing a metric that matches what is actually cared about, rather than whichever one the library prints first. Calibration and thresholds: how a model’s score turns into a decision about a person. Prediction against explanation, and when each is wanted.
Activity: students move the threshold on a classifier and see who changes category.
Suggested reading: The Art of Statistics, ch. 6; Google’s free notes on overfitting and generalisation, classification metrics, and thresholds.
Decision trees and random forests, conceptually rather than mathematically. Classification. Clustering, for when nothing is labelled. The running theme: AI is not magic, and garbage in still means garbage out.
Hands-on: build a classifier, compare it against a deliberately weak baseline, and find out how often the weak baseline wins.
Suggested reading: ISLP, ch. 2, concepts only; official chapter 2 slides; Google’s clustering overview.
Taught for the first hour: data leakage and shortcut learning, where a model is effectively cheating but its accuracy score looks excellent. Dataset shift, and why models that work in testing fail once deployed. Fairness basics, and checking whether a model works as well for one group as for another.
Hands-on for the second hour: students are given a model with suspiciously good numbers and have to find where it is cheating.
Quiz 4: modelling, evaluation and trustworthiness.
Suggested readings: AI Snake Oil, ch. 2–3; Princeton’s free excerpt covers predictive AI; Kapoor and Narayanan, “Leakage and the Reproducibility Crisis in ML-based Science” (2023, open access).
From regression to neural networks, conceptually. What an LLM does: next-token prediction, training data, and where the unexpected capabilities come from. Why the tool that has been writing the class’s code all term can be completely fluent and completely wrong at once. Hallucination, bias, and the limits of pattern-matching.
Project reports are due at the start of this session.
Suggested readings: Stephen Wolfram, “What Is ChatGPT Doing … and Why Does It Work?”; Jay Alammar, “The Illustrated Transformer”.
This session comes late in the course deliberately: by now the class has used AI on a complete project and knows enough to judge what comes back. Using LLMs to clean data, generate code, summarise findings and suggest analyses. Prompt patterns that work for data work. When to check the output, and how. Assembling a workflow worth using again.
Run as a retrospective: every student brings the worst thing ChatGPT did to them during the project, and the class works out how it should have been caught.
Algorithmic bias and fairness, using Indian and international cases. Privacy and consent now that collecting data costs nothing. Who is accountable when a system makes the decision: hiring, lending, healthcare, criminal justice. Then the communication half: building an argument from data as claim, evidence and warrant, and writing about data for readers who will never see the code.
The session closes the course by returning to the three questions it opened with.
Quiz 5: LLMs, ethics and communication.
Suggested readings: Weapons of Math Destruction, introduction and ch. 1, 5, 6; Cathy O’Neil’s shorter interview on the same argument; Data Feminism, ch. 1 and ch. 7; Angwin et al., “Machine Bias”, ProPublica (2016). For the Indian material: a chapter from Reetika Khera (ed.), Dissent on Aadhaar (Orient BlackSwan, 2019), and the consent and purpose-limitation provisions of the Digital Personal Data Protection Act, 2023.
Held free. Normally used to debrief the examination and to discuss where students might go next, including the Data Science minor.