Skip to main content

Hephaestus — DPIA Pre-Screen (Art. 35 GDPR)

Status: controller/DPO determination pending. Several WP29 risk criteria are present. Advisory output and the absence of grading, employment, or access decisions reduce impact but do not negate evaluation, systematic monitoring, dataset combination, or the student/employee power imbalance. This engineering screen is not the controller's Art. 35 determination.

Review owner: TUM/AET data-protection coordinator. Engineering review: 2026-08-04. Next review: before first production deployment or 2026-11-04, whichever comes first. Controller decision reference: pending.

Records whether a full Data Protection Impact Assessment is required for the TUM-operated Hephaestus deployment. Two gates apply to TUM as a Bavarian public body: Art. 35(3) GDPR and the Bavarian Blacklist published by the BayLfD under Art. 35(4) GDPR. The DSK list (which addresses the non-public sector) is referenced only as a cross-check.

1. Threshold check against Art. 35(3) GDPR

TriggerPresent?Reasoning
Systematic and extensive evaluation, including profiling, that forms the basis for decisions producing legal effects or similarly significant effects (Art. 35(3)(a))NoHephaestus produces advisory observations and guidance delivered to the contributor and visible to other workspace members on dashboards and per-artefact views. They are not consumed by any automated grading, assessment, HR, or access-control pipeline operated by Hephaestus. The processing is workspace-scoped and practice-scoped (defined per workspace by the workspace administrator), not exhaustively profiling individuals across a large population.
Large-scale processing of Art. 9(1) (special categories) or Art. 10 (criminal convictions) dataNoHephaestus does not intentionally solicit or classify these data. Free-text repository content may contain and therefore cause incidental processing of them, but processing is limited to explicitly connected repositories and is not systematic or large-scale.
Systematic monitoring of a publicly accessible area on a large scaleNoConnected repositories may be public, but Hephaestus does not crawl or discover public sources. It processes only repositories explicitly connected by a workspace administrator, at the current deployment scale.

None of the three Art. 35(3) triggers is present.

2. Bavarian Blacklist (BayLfD)

The BayLfD's Bavarian Blacklist (published 7 March 2019 under Art. 35(4) GDPR) enumerates concrete public-sector processing constellations that require a DPIA. None of the listed constellations describes a university teaching-support feedback platform. The risk-criteria assessment below tracks the WP29 Guidelines on DPIA (WP248rev.01) criteria the BayLfD applies in screening:

WP29 criterionPresent?Current facts and safeguards
Evaluation or scoringYesPractice reviews create contributor-specific assessments of observed engineering work. Advisory and contestable output, no automated grading/HR/access decision, and separate delivery controls reduce consequence; they do not negate the criterion.
Automated decision with legal or similarly significant effectNoObservations do not determine grades, employment, recognition caps, feature access, merge rights, or another legal/significant outcome. Any such consumer is a material change requiring a full reassessment before use.
Systematic monitoringYesEnabled workspaces repeatedly observe activity in administrator-connected repositories and optional monitored Slack/Outline sources. The scope is selected rather than an open crawl, but the processing is organized and recurring.
Sensitive or highly personal dataPartlyThe system does not solicit or classify Art. 9/10 data, but free-text code, issues, messages, documents, and conversations can contain incidental sensitive or highly personal content. Private mentor and Slack conversations receive the stricter source class.
Data processed on a large scaleNo at the documented deployment scaleThe deployment is workspace-scoped and currently limited. The controller must reassess before a material population, repository, retention, or source expansion.
Matching or combining datasetsYesDetection and mentoring can combine GitHub/GitLab activity with Hephaestus observations and, where enabled, selected Slack messages and Outline documents. This is source combination even though Hephaestus does not enrich people from commercial external profiles.
Vulnerable data subjectsYesStudents and employees can face a power imbalance and may be unable to oppose workspace-level processing as easily as an ordinary consumer. Advisory use, rights procedures, and the ban on grading/HR use are mitigations.
Innovative technology or organizational solutionYesAn LLM interprets work artifacts and produces contributor-specific practice feedback. Enterprise no-training terms and restricted egress reduce risk but do not remove this criterion.
Prevents exercise of a right or use of a service/contractNoThe product does not gate course, employment, repository, or Hephaestus access on observations.

WP248 rev.01 states that, in most cases, processing meeting two criteria warrants a DPIA and that the more criteria are present, the more likely high risk becomes. Multiple criteria above are present. The controller/DPO must therefore record either a full DPIA or a reasoned determination that one is not required.

3. DSK list — cross-check (not directly applicable)

The DSK list of processing operations requiring a DPIA under Art. 35(4) GDPR addresses the non-public sector and is not directly applicable to TUM/AET as a Bavarian public body; the controlling instrument for TUM is the Bavarian Blacklist in §2. The absence of one exactly named DSK constellation does not override the multi-criterion WP29 assessment above. Enterprise no-training terms, the documented joint-controller arrangement, restricted LLM egress, and the Art. 21 rights process reduce residual risk but do not decide whether Art. 35 requires a DPIA. They remain mandatory while the determination is pending.

4. Residual risk analysis

RiskLikelihoodImpactMitigation
Free-text artefact content (PR descriptions, issue descriptions, commit messages, review comments) contains personal data of identifiable third parties and is transmitted to the LLM providerLow-mediumLow-mediumThe privacy statement instructs users not to enter third-party personal data. The practice-review sandbox runs on a per-job --internal Docker network with no general egress except a token-authenticated LLM proxy, and the provider is subject to enterprise no-training and transfer safeguards. Objections to processing follow the Art. 21 contact process in privacy §7.
Practice-feedback comments or Slack reminders reach a contributor after their comments-and-reminders setting has been disabled, or the Slack mentor accepts a new interaction after the workspace mentor has been disabledLowLow-mediumDelivery policy is checked before starting each comment or reminder and before admitting each new Slack mentor interaction. An interaction admitted before the workspace switch changes may finish. The personal setting covers issue, pull-request, and merge-request comments plus Slack reminders, and delivered comments link back to it.
LLM provider retains the prompt beyond the enterprise default retention windowLowLow-mediumEnterprise no-training terms; abuse-monitoring retention per the provider's published terms (Microsoft Azure OpenAI: enterprise abuse-monitoring window per the published data-privacy documentation; eligible customers can apply for Microsoft's modified abuse monitoring / Limited Access program); Zero Data Retention can be negotiated where the provider supports it; DPF / SCCs Module 2 in place.
Server access logs retain IP addressesLowLowNo HTTP access log is written at any layer: Tomcat's is explicitly disabled in the production profile, the Traefik reverse proxy is not started with --accesslog (Traefik's default is off), and both nginx containers disable it at the server level, so no per-request IP/URL record is created. IP addresses are recorded only against authentication events, under that log's 12-month partitioned window.
gitlab.lrz.de content leaks via the LRZ integrationLowLowLRZ is a separate controller; inter-public-body transmission under Art. 5(1) Nr. 1 BayDSG; LRZ applies its own TOMs on its own infrastructure.
Workspace administrator enables Slack without informing contributors about monitored channelsLow-mediumLowJoint-controller / shared-responsibility model documented in privacy §10; monitored channels are forward-only, require explicit activation, post a visible channel announcement, and provide App Home/settings opt-out plus erasure.
Compromise of the application DB exposes federated identity links + cookie-session revocation listLowMediumSelf-hosted on AET infrastructure; upstream tokens encrypted at rest; short-lived ES256 cookie-JWTs with server-side revocation; TLS-only ingress; incident response under TUM DPO oversight.
Source expansion combines more contributor context than an enabled practice needsMediumMedium-highResolve the exact practice set before collection; compile a minimum evidence plan; default-deny every new source/purpose; require the artifact-source governance gate.
A missing source systematically withholds feedback from particular platforms, workflows, or privacy choicesMediumMediumThe internal readiness report records refusal separately from what a review observed. Do not compare or rank results across unequal evidence coverage. Coverage analytics and administrator remediation must not be claimed until their operator surface is implemented.
Repository history exposes deleted secrets or personal data beyond the reviewed changeMediumHighDefault review bundles use bounded immutable .git-free snapshots. Repository history is a separate governed source and is not implied by tree access.
Restricted or reviewer-only context is quoted into learner-facing feedbackLow-mediumHighSource-use decisions govern automated review and feedback delivery separately. Capture requires the automated-review purpose; delivery rechecks the feedback-delivery purpose, source authorization, and citation ownership.

5. Safeguards that must remain in place

  • No-training enterprise API terms for every LLM provider configured by the AET-pool. Regressing to a consumer tier is a material change.
  • Per-job LLM proxy enforced by the practice-review sandbox. DNS disabled inside the sandbox; outbound traffic limited to a per-job, token-authenticated proxy. Any widening of this network posture is a material change.
  • Delivery controls: a signed-in contributor's Comments and Slack reminders setting and the workspace mentor switch gate their corresponding delivery paths. Removing or bypassing either is a material change.
  • Data-subject rights process in privacy §7. The delivery setting is not presented as an Art. 21 objection or as a control over review processing.
  • Workspace-administrator joint-controller notice in the privacy statement (§10). Structural changes to the shared-responsibility split require an amended record.
  • No HTTP access log. Tomcat's access log is disabled in the production profile, the Traefik ingress runs without --accesslog, and both nginx containers disable it at the server level, so request-level IP/URL data is never written at any layer. Enabling it at any layer is a material change.
  • Error telemetry and product analytics remain disabled. The webapp ships a Sentry integration and a PostHog integration, both disabled in the current production deployment. Activating either is a material change.

6. Required determination and change freeze

None of the specific Art. 35(3) examples is documented as present, but the broader WP29 screen identifies several concurrent risk criteria. The controller must record the Art. 35 determination before processing expands.

Until that decision is recorded, the TUM-operated deployment must not enable a new artifact-source family, a new cross-source purpose, broader processor egress, extended evidence retention, or a materially broader monitored population. This is not an instruction to weaken existing safeguards or erase operational audit evidence.

ENGINEERING_BASELINE / ENGINEERING_APPROVED in the machine source-use registry is not a legal decision and does not permit expansion.

Regardless of the pending determination, a DPIA must be opened or amended before any of the following takes effect:

  • LLM provider is added, changed to a consumer tier, or loses its no-training commitment.
  • Practice-review sandbox gains outbound connectivity beyond the per-job LLM proxy.
  • Observations begin to drive any automated decision within Hephaestus (grading, recognition caps, feature access).
  • The bundled Sentry integration is activated against a SaaS tenant, or the bundled PostHog integration is activated.
  • The processing population starts to include data subjects in a category covered by the BayLfD vulnerable-data-subjects criterion.
  • Repository ingestion expands beyond administrator-selected repositories into systematic or large-scale monitoring of public sources.
  • A new artifact source, source combination, private-conversation use, repository-history use, research/evaluation reuse, retention extension, or learner/admin audience is proposed and is not already covered by the recorded decision.

The source-specific decision and test checklist lives in artifact-source-governance.md. The controller's decision identifier and date must replace the pending status at the top of this file when the determination is recorded.