Where the numbers come from
- Grades: MadGrades, which republishes UW–Madison's public grade reports by term, section and instructor. GPA counts letter grades only (A=4, AB=3.5, B=3, BC=2.5, C=2, D=1, F=0); S/U, credit and incomplete grades are left out.
- Catalog and current sections: the public Course Search & Enroll API. Credits, breadths, prerequisites and who teaches this term come from there. Seat numbers are a snapshot from the build date, so check Enroll before you plan around them.
- Chat mentions: how often a course code came up in a UW–Madison student group chat. Only aggregate counts were ever exported: no messages, names or ids.
- Reddit: public threads in r/UWMadison that name the course, found by web search. Only the thread title and link are stored. Tone is a rough keyword score (words like "easy" vs "cooked").
- Rate My Professors: not copied. Each instructor has a link that searches RMP at UW–Madison.
How courses were checked
Course codes pulled from chat are messy ("cs300", "comp sci 300", "stats 240"). Each one was matched against the real catalog, and codes that don't exist were dropped, including a few that were really "comp sci" misread as nutrition science. Cross-listed courses such as COMP SCI 240 / MATH 240 were merged into one row. The largest courses of the current term were added even if nobody mentioned them. Every other course in the catalog is listed too, with grades and seats but no chat or Reddit data.
Reading the derived numbers
- Recent GPA uses the last six terms with grades.
- Trend is a weighted least-squares slope of term GPA since fall 2016, in GPA points per year.
- Instructor spread is the gap between the most and least generous instructor with at least 30 graded students in that course.
- Adjusted GPA (used for "Highest GPA") pulls small courses toward their subject's average, so a 12-student seminar with a 4.0 doesn't outrank a 900-student course with a 3.9.
Open data
The full SQLite database (data/uwcourses.db) and the scripts that build it are in the GitHub repo. The analysis write-up lives in docs/ANALYSIS.md.
数据从哪来
- 成绩:来自 MadGrades,它按学期、section 和老师公开 UW–Madison 的官方成绩报告。GPA 只算字母成绩(A=4、AB=3.5、B=3、BC=2.5、C=2、D=1、F=0),S/U、学分制和 Incomplete 不计入。
- 课程目录和本学期开课:来自公开的 Course Search & Enroll 接口,学分、breadth、先修要求和本学期老师都从这里拿。座位数是建站当天的快照,选课前请以选课系统为准。
- 群聊提及:某个课号在一个 UW–Madison 学生群里出现的次数。导出的只有汇总计数,没有任何消息内容、名字或 ID。
- Reddit:r/UWMadison 里提到这门课的公开帖子,通过网页搜索找到。只保存帖子标题和链接。语气是粗略的关键词打分(比如 "easy" 和 "cooked")。
- Rate My Professors:没有复制任何内容,每位老师旁边只放了一个跳到 RMP 搜索的链接。
课程是怎么核对的
群聊里的课号写法很乱("cs300"、"comp sci 300"、"stats 240")。每个课号都和真实的选课目录比对过,不存在的直接删掉,其中有几个其实是 "comp sci" 被误认成了营养学。交叉挂牌的课(比如 COMP SCI 240 / MATH 240)合并成一行。本学期人数最多的大课即使没人提过也一并收录;目录里的其余课程也都列出,有成绩和座位,但没有群聊和 Reddit 数据。
衍生指标怎么看
- 近 6 学期 GPA:只算最近六个有成绩的学期。
- 趋势:2016 年秋以来每学期 GPA 的加权最小二乘斜率,单位是每年多少绩点。
- 老师差距:同一门课里给分最松和最严的两位老师之间的 GPA 差(每位至少 30 个学生)。
- 调整后 GPA("GPA 最高"排序用的就是它):把小班课往本学科平均值拉,免得 12 人的研讨课靠运气拿 4.0 排到 900 人、3.9 的大课前面。
开放数据
完整的 SQLite 数据库(data/uwcourses.db)和生成脚本都在 GitHub 仓库里,分析报告在 docs/ANALYSIS.md。