Privacy Policy

This page describes how DS593's submission system collects, uses, shares, and retains your data.

Purpose of collection

The data described on this page is collected solely to run the course: to distribute lecture questions, and to collect and grade your answers.

Data we collect

The system stores the following personal data about you:

The system also stores course content that is not personal data: the list of lectures and their questions.

Who can access your data

Consent to improve the course

Separately from the above, you are asked to consent to automated analysis of your answers to improve the course. This is optional, and your choice never factors into grading or course performance in any way.

If you consent:

You can change your consent choice at any time by returning to the signup page and submitting the form again, with the consent box checked or unchecked as you prefer — this updates your consent without changing your API key or any other data. Changing your choice is not retroactive: any analysis already carried out while you were consenting cannot be undone.

Subject access rights

You may access a copy of every piece of personal data this system holds about you at any time, on the Access my data page. That page is generated directly from the tables below, so it always reflects exactly what is stored.

You may not request deletion of your data until the retention period described below has elapsed.

Data retention

Your data is retained for the duration of the course, and for up to one year after final grades are submitted, to allow for grade disputes and ordinary academic record-keeping. Boston University's published Record Retention Table does not list a specific entry for raw student coursework submissions themselves; this course's one-year window is chosen to be consistent with the retention periods it does specify for closely related grading records (e.g. professors' grade books).

Data security and third parties

Your email and answers are educational records protected under FERPA, and are classified as Confidential under BU's Data Classification Policy. In accordance with that policy, this data is stored and processed only on University-approved infrastructure, and is only shared with a third-party vendor or service (including any external language model used for the automated analysis above) once that vendor has been approved by BU Information Security and meets BU's applicable security standards.


Technical Details

A few words about how we use some papers and systems from our course to enforce and realize this privacy policy.

Implementing GDPR Subject Access Requests with K9db

We annotate the schema below using the data-ownership annotations K9db introduces, to show how your data is protected:

CREATE DATA_SUBJECT TABLE users (
    email varchar(255),
    apikey varchar(255),
    is_admin int,
    consent int,
    PRIMARY KEY (email),
    UNIQUE (apikey)
);

-- lectures and questions table are unsharded.
CREATE TABLE lectures (
    id int,
    label varchar(255),
    PRIMARY KEY (id)
);
-- question_number: number *within* lecture
CREATE TABLE questions (
    id int AUTO_INCREMENT,
    lecture_id int,
    question_number int,
    question text,
    PRIMARY KEY (id),
    FOREIGN KEY (lecture_id) REFERENCES lectures(id)
);

-- Answers are owned by the student that provided the answer.
CREATE TABLE answers (
    id varchar(255),
    email varchar(255),
    lec int,
    question_id int,
    answer text,
    submitted_at datetime,
    PRIMARY KEY (id),
    FOREIGN KEY (email) OWNED_BY users(email),
    FOREIGN KEY (lec) REFERENCES lectures(id),
    FOREIGN KEY (question_id) REFERENCES questions(id)
);

-- A presenter owns the record that marks them as a presenter of some lecture.
CREATE TABLE presenters (
    id int AUTO_INCREMENT,
    lecture_id int,
    email varchar(255) OWNED_BY users(email),
    PRIMARY KEY (id),
    FOREIGN KEY (lecture_id) REFERENCES lectures(id)
);

Enforcing Data Use Policies Using Sesame

The rules described earlier on this page — who can read your email, your answers, or your API key, and under what circumstances — are not just written down; we encode them as code and enforce them automatically using Sesame. This means the policies remain upheld even if we introduced a bug into the application.

We attach Sesame policies to each database column that holds personal data, and every time the application attempts to read that column, Sesame attaches the corresponding policy to the application's read data.

Whenever the application tries to externalize some data to the outside world, e.g., by showing it in an HTML page, Sesame automatically checks all policies associated with that data. Sesame refuses to externalize the data if the policy check fails.

We have defined the following policy:

Simplified, these checks read like this:

// QueryableOnly, on users.apikey
fn check(reason) -> bool {
    match reason {
        Reason::Query(sql) => sql.starts_with("SELECT"),
        Reason::CookieSet("apikey") => true,
        _ => false, // rendering it on a page, or anything else: refused
    }
}

// AnswerAccessPolicy, on answers.lec/question_id/answer/submitted_at
fn check(who_is_asking) -> bool {
    if who_is_asking == self.owner { return true }         // you wrote it
    if who_is_asking.is_admin() { return true }            // the instructor
    if presenters_of(self.lecture).contains(who_is_asking) {
        return true                                         // you present this lecture
    }
    false
}

// AutomatedAnalysisPolicy, on consented_answers.answer
fn check(who_is_asking, purpose) -> bool {
    if !self.consent { return false }                  // you did not consent
    if purpose != AutomatedAnalysisRequest { return false } // only for that specific use
    who_is_asking.is_admin()                           // only the instructor can trigger it
}

These are simplified for readability; see the Sesame paper for the full policy framework.