AI and Data Analytics

Academic Programme UG B.A. (Hons) / B.Sc. (Hons)
Academic Year / Term(s) 2026-27 / Term 4
Course Credits 4
Course Type -
Course Category Mandatory
Host Discipline -
Cross-listed Discipline(s) -
Pre-requisites CS 101 Principles of Computer Science and Software Development; Mathematics and Logic (Part 2, Introduction to Statistics)
Instructor(s) Navin Kabra
Instructor Email Id(s) navin@smriti.com
Faculty Associate (FA) Devdatta Kathale; Ashvina Kulkarni
FA Email Id Devdatta.Kathale@nayanta.edu.in; Ashvina.Kulkarni@nayanta.edu.in

Course Description

This course teaches students to think carefully about data and AI: to ask the right questions, find or collect the right evidence, choose an appropriate analysis, and report what they find honestly.

The focus is on three questions:

  1. Given some data, what analysis can and should I do on it?
  2. What analysis is justified and supported by the data?
  3. Given some analysis I want to do, what kind of data do I need?

AI appears in three ways. As a subject of study: how do these systems learn from data, and where do they fail? As a programming tool: students use an LLM to write their code throughout the course. And as an analysis tool: to be able to quickly do many different analyses, to quickly see the consequences of different choices, and in general do a broader exploration of the design space.

Students will have a ChatGPT Plus subscription for the duration of the course, and will use it in almost every session. No prior experience with pandas, matplotlib or machine learning is assumed.

Learning Objectives

The course aims to give students:

Learning Outcomes

By the end of this course, students will be able to:

Teaching Methods / Pedagogy

The course runs over six weeks, in eighteen sessions: twelve of 90 minutes and six of two hours. A seventh week is used for assessment.

Each week has two 90-minute Tuesday sessions and one two-hour Thursday session. The Tuesday sessions introduce material, and the second one is always hands-on rather than a second lecture. Many of the Thursday sessions are labs or studios.

Specific choices worth noting:

Course Schedule

S.No. Topics Suggested Readings
1 The Data Mindset: questions, evidence and reasoning Calling Bullshit, ch. 1–2; free course videos
2 Exploring data with AI-assisted code —
3 Lab: environment, and one dataset end to end —
4 When summary statistics mislead The Art of Statistics, ch. 1–2
5 Visualisation as argument How Charts Lie, ch. 1–3; Data Feminism, ch. 3; Cairo’s short guide
6 Studio: duelling charts Calling Bullshit, ch. 7; free visualisation notes
7 Where data comes from Calling Bullshit, ch. 4; Data Feminism, ch. 6
8 Correlation, causation and study design The Art of Statistics, ch. 4; free causality videos
9 Lab: data you collect yourself, including data that isn’t a table —
10 Hypothesis testing, and the replication crisis The Art of Statistics, ch. 10–11; Ioannidis (2005)
11 What is a model? From regression to AI The Art of Statistics, ch. 5; ISLP ch. 2 slides
12 Lab: the p-hacking sandbox Gelman & Loken, “The Garden of Forking Paths”; FiveThirtyEight’s article and interactive
13 Evaluation: generalisation, metrics and decisions The Art of Statistics, ch. 6; Google’s notes on overfitting and classification metrics
14 Trees, forests and clustering ISLP ch. 2 (concepts only); official ch. 2 slides; Google’s clustering overview
15 Trustworthy AI: leakage, shift and failure modes; leakage hunt AI Snake Oil, ch. 2–3; Princeton excerpt; Kapoor & Narayanan on leakage
16 Inside the AI: how large language models work Wolfram, “What Is ChatGPT Doing?”; Alammar, “The Illustrated Transformer”
17 AI as your data analysis partner —
18 Ethics, consequences and communicating data Weapons of Math Destruction, ch. 1, 5, 6; author interview
Assessment day: final examination and project defence —
Free slot: Q&A and debrief —

Nothing on this list is required. The list is for students who want to follow something up, or who learn better from books and articles.

Two books cover much of the course:

These books/articles will be useful for particular parts of the course:

References for looking things up rather than reading:

Shorter pieces, all free, are listed against the individual sessions below.

There is no Python textbook on this list, and that is intentional. Students have ChatGPT Plus, and the course teaches them to direct it and check its output rather than to memorise syntax.

Assessment Methods

S.No. Component Weightage
1 Quizzes: five short in-class MCQ quizzes, best four counted 20%
2 Project, in pairs: proposal, analysis and written report along with all data and code 30%
3 Project defence: individual, written, 60 minutes 25%
4 Final examination: written, 2 hours 25%

Quizzes. Five quizzes of about ten minutes, at the end of a session, from Week 2 onwards. Each tests whether a concept from recent classes were learnt properly. The best four of five count, so one missed quiz costs nothing.

Project. To be done in pairs. Projects will be self-selected with the help and approval of the instructor and the faculty associates. All projects should involve an important question, some data curation and cleanup, followed by analysis, and a written report, along with all the data and the programs. The report must explain how and where AI was used in the project, and how the outputs of the AI were verified.

Project defence. After the project report is submitted and assessed by the FAs and instructor, each student receives a printed sheet of five questions written specifically about their own project, and answers them by hand in an invigilated room. Students may have their own report on the desk and nothing else. The two members of a pair get different questions and answer individually and independently.

Final examination. A two hours written, invigilated exam. Questions will typically contain some data artifact — a chart, a regression output, a description of how a sample was collected, an extract of AI-generated analysis — and will ask the student to judge it. This test will not require calculations.

Class Policies

Attendance

There are no marks for attendance or classroom participation.

However, keep in mind that the 5 quizzes happen during class, and will not be repeated. Students who miss two or more sessions with quizzes could lose marks.

In addition, university attendance requirements apply as normal.

Academic Integrity, Plagiarism, and AI Usage

This course requires the use of AI. Use of AI in all aspects of the course is encouraged. Every student has a ChatGPT Plus subscription paid for by the university, and most of the code written during the course will be produced by an LLM rather than typed by hand.

However, to ensure that students actually learn things while using AI, a major part of the evaluation of the course (quizzes, project defense, and final exam) will be offline, without the use of AI or any devices.

Verification. Students are responsible for everything they submit, including the parts they did not write. “ChatGPT produced that” is not a defence for a wrong number, a misleading chart, or an analysis that does not answer the question.

Passing off another student’s work, or published text, as one’s own is dealt with under the university’s normal plagiarism policy. This is unaffected by the AI provisions above.

Requirements / Provisions

EXTENDED COURSE OUTLINE

Sessions are grouped by teaching week. Module labels are given in brackets. All readings listed are suggested; none is required.

Week 1 — Reasoning with data, and first contact

Session 1: The Data Mindset — Questions, Evidence, and Reasoning (Module 1)

What data analytics is, how it is different from statistics, and why it is worth studying. The cycle: question, data, analysis, interpretation, communication. Types of data — structured, unstructured, qualitative, quantitative. How to take a vague question and make it precise enough to answer with data. Confounding, bias, and the limits of what can be learned from watching rather than intervening.

Activity: students will be given some recent newspaper headlines and work out what data and analysis would support it, and what would not.

Suggested reading: Calling Bullshit, ch. 1–2. The course videos cover the same introductory material for free.

Session 2: Exploring Data with AI-Assisted Code (Module 2)

Introduction to pandas and matplotlib, with ChatGPT writing the code. Loading, inspecting and summarising a dataset. Students write the prompt, then read the code that comes back and check whether it does what they asked for. This session establishes the working method for the rest of the course.

Session 3: Lab — Environment, and One Dataset End to End (Module 2)

Getting every student’s Python and ChatGPT Plus working. Then one small dataset all the way through: load it, look at it, summarise it, produce one chart, write one sentence about what it shows. None of it is difficult; the point is that all thirty students have been through the complete loop once.

An ungraded diagnostic quiz at the end, to establish what the class remembers from CS 101 and from the statistics half-course.

Week 2 — Summaries, and charts as argument

Session 4: When Summary Statistics Mislead (Module 2)

Connecting the statistics the class already knows to real, messy data. Why the mean is often the wrong number to report: skew, outliers, Simpson’s paradox. Choosing a summary that answers the question being asked.

Suggested reading: The Art of Statistics, ch. 1–2.

Session 5: Visualisation as Argument (Module 2)

The same data can be charted to tell very different stories. Tufte’s principles of chart design. The usual ways charts mislead: truncated axes, cherry-picked date ranges, aggregation that hides the interesting variation. What ChatGPT produces by default when asked for “a chart”, and what that default assumes on the analyst’s behalf.

Suggested readings: How Charts Lie, ch. 1–3; Data Feminism, ch. 3, “On Rational, Scientific, Objective Viewpoints”; Wilke, Fundamentals of Data Visualization, ch. 5. Cairo covers much of the same ground in this shorter illustrated article.

Session 6: Studio — Duelling Charts (Module 2)

One dataset, class split in two. Each half is assigned a conclusion and has to build the best honest chart it can in support of it. Both sides present, and the class then argues about where the line falls between framing something well and misleading people.

Quiz 1: summaries and visualisation.

Suggested reading: Calling Bullshit, ch. 7. The free visualisation notes cover misleading axes and proportional ink.

Week 3 — Where data comes from

Session 7: Where Data Comes From (Module 3)

Sampling, surveys, sensors, scraped web pages, administrative records. Selection bias, survivorship bias, measurement error. The difference between data collected deliberately and data that was found.

Case study: biased training data in deployed AI systems, for example in predictive policing and in medical imaging.

Suggested readings: Calling Bullshit, ch. 4; Data Feminism, ch. 6, “The Numbers Don’t Speak for Themselves” (complete chapter online).

Session 8: Correlation, Causation, and Study Design (Module 3)

Observational against experimental data. Why correlation is not causation, with enough examples that it sticks. Randomised controlled trials, natural experiments, A/B testing. Then the reverse question: given a causal question, what data would be needed to answer it? And why a model trained on observational data cannot answer it, however accurate the model looks.

Suggested reading: The Art of Statistics, ch. 4; the Calling Bullshit causality videos cover the same ideas for free.

Session 9: Lab — Data You Collect Yourself, Including Data That Isn’t a Table (Module 3)

Designing a survey or collection instrument, and the many ways it goes wrong. Then the harder case: images, audio, video, sensor readings. What is involved in turning any of those into something that can be computed with, and what gets thrown away in the process.

The project brief is issued in this session and pairs are formed.

Quiz 2: provenance and causation.

Week 4 — Inference, and the first model

Session 10: Hypothesis Testing, and the Replication Crisis (Module 3)

p-values and confidence intervals applied to real data. What a p-value does and does not permit a student to claim. Effect size against statistical significance. Then the replication crisis in psychology and medicine: multiple testing, the garden of forking paths, publication bias, and how all of it happens without anyone setting out to cheat.

Suggested readings: The Art of Statistics, ch. 10–11; Ioannidis, “Why Most Published Research Findings Are False”, PLoS Medicine (2005, open access).

Session 11: What Is a Model? From Regression to AI (Module 4)

A model is a simplified description of the world, learned from data. Linear regression as the simplest example of something that learns: fitting a line, interpreting the coefficients, looking at the residuals. Why a baseline is needed before any modelling result means anything.

Suggested reading: The Art of Statistics, ch. 5; ISLP’s chapter 2 slides are a free introduction to models, features, targets, prediction and inference.

Session 12: Lab — The p-hacking Sandbox (Module 3)

Students are given a dataset of pure noise and asked to find a publishable result in it. Most will manage it. They then write down an analysis plan for a real dataset before looking at it, and hold themselves to it.

Project proposals are due in this session, with the analysis plan written down in advance, and each receives feedback in class.

Quiz 3: inference.

Suggested readings: Gelman and Loken, “The Garden of Forking Paths”; FiveThirtyEight, “Science Isn’t Broken”, including its p-hacking tool.

Week 5 — Models, and whether to trust them

Session 13: Evaluation — Generalisation, Metrics, and Decisions (Module 4)

Train/test splits and cross-validation, at the level of intuition rather than formulae. Overfitting and underfitting. Choosing a metric that matches what is actually cared about, rather than whichever one the library prints first. Calibration and thresholds: how a model’s score turns into a decision about a person. Prediction against explanation, and when each is wanted.

Activity: students move the threshold on a classifier and see who changes category.

Suggested reading: The Art of Statistics, ch. 6; Google’s free notes on overfitting and generalisation, classification metrics, and thresholds.

Session 14: Trees, Forests, and Clustering (Module 4)

Decision trees and random forests, conceptually rather than mathematically. Classification. Clustering, for when nothing is labelled. The running theme: AI is not magic, and garbage in still means garbage out.

Hands-on: build a classifier, compare it against a deliberately weak baseline, and find out how often the weak baseline wins.

Suggested reading: ISLP, ch. 2, concepts only; official chapter 2 slides; Google’s clustering overview.

Session 15: Trustworthy AI — Leakage, Shift, Failure Modes; and a Leakage Hunt (Module 4)

Taught for the first hour: data leakage and shortcut learning, where a model is effectively cheating but its accuracy score looks excellent. Dataset shift, and why models that work in testing fail once deployed. Fairness basics, and checking whether a model works as well for one group as for another.

Hands-on for the second hour: students are given a model with suspiciously good numbers and have to find where it is cheating.

Quiz 4: modelling, evaluation and trustworthiness.

Suggested readings: AI Snake Oil, ch. 2–3; Princeton’s free excerpt covers predictive AI; Kapoor and Narayanan, “Leakage and the Reproducibility Crisis in ML-based Science” (2023, open access).

Week 6 — Inside AI, ethics, communication

Session 16: Inside the AI — How Large Language Models Work (Module 5)

From regression to neural networks, conceptually. What an LLM does: next-token prediction, training data, and where the unexpected capabilities come from. Why the tool that has been writing the class’s code all term can be completely fluent and completely wrong at once. Hallucination, bias, and the limits of pattern-matching.

Project reports are due at the start of this session.

Suggested readings: Stephen Wolfram, “What Is ChatGPT Doing … and Why Does It Work?”; Jay Alammar, “The Illustrated Transformer”.

Session 17: AI as Your Data Analysis Partner (Module 5)

This session comes late in the course deliberately: by now the class has used AI on a complete project and knows enough to judge what comes back. Using LLMs to clean data, generate code, summarise findings and suggest analyses. Prompt patterns that work for data work. When to check the output, and how. Assembling a workflow worth using again.

Run as a retrospective: every student brings the worst thing ChatGPT did to them during the project, and the class works out how it should have been caught.

Session 18: Ethics, Consequences, and Communicating Data (Module 6)

Algorithmic bias and fairness, using Indian and international cases. Privacy and consent now that collecting data costs nothing. Who is accountable when a system makes the decision: hiring, lending, healthcare, criminal justice. Then the communication half: building an argument from data as claim, evidence and warrant, and writing about data for readers who will never see the code.

The session closes the course by returning to the three questions it opened with.

Quiz 5: LLMs, ethics and communication.

Suggested readings: Weapons of Math Destruction, introduction and ch. 1, 5, 6; Cathy O’Neil’s shorter interview on the same argument; Data Feminism, ch. 1 and ch. 7; Angwin et al., “Machine Bias”, ProPublica (2016). For the Indian material: a chapter from Reetika Khera (ed.), Dissent on Aadhaar (Orient BlackSwan, 2019), and the consent and purpose-limitation provisions of the Digital Personal Data Protection Act, 2023.

Week 7 — Assessment

Assessment day (Tuesday)

Free slot (Thursday)

Held free. Normally used to debrief the examination and to discuss where students might go next, including the Data Science minor.