Skip to content

Data Protection Impact Assessment (DPIA) ​

SpektraBot SEND Support System ​

FieldDetail
Document version1.0
Date[SIGN-OFF REQUIRED]
Author[SIGN-OFF REQUIRED]
Data ControllerSpectrum Dynamics CIC
DPO / Privacy Lead[SIGN-OFF REQUIRED]
System nameSpektraBot
ICO registration number[SIGN-OFF REQUIRED]

Document Control ​

VersionDateAuthorChanges
1.0[SIGN-OFF REQUIRED][SIGN-OFF REQUIRED]Initial DPIA

Step 1: Identify the Need for a DPIA ​

1.1 Why is a DPIA required? ​

UK GDPR Article 35 requires a Data Protection Impact Assessment when processing is likely to result in a high risk to the rights and freedoms of natural persons. The ICO's screening criteria mandate a DPIA when two or more of the following apply. SpektraBot meets six:

ICO Screening CriterionApplicable?Justification
Evaluation or scoringYesAI-generated advice based on user-provided child profile data; conversation intent detection; document relevance scoring
Automated decision-making with legal or significant effectsYesAI-generated guidance on EHCP applications, tribunal appeals, and legal rights directly affects families' educational provision decisions
Systematic monitoringNoSpektraBot does not systematically monitor public areas
Sensitive data or data of a highly personal natureYesSEND diagnoses, EHCP status, therapy details, mental health indicators (crisis detection), and safeguarding concerns
Data processed on a large scaleYesDesigned to serve families across 152 English local authorities; knowledge base of 450K+ document chunks across all LAs
Datasets that have been matched or combinedYesCombines user-provided child profiles with GIAS school data (52K schools), LA EHCP statistics, Ofsted ratings, crawled Local Offer content, and SEND tribunal data
Data concerning vulnerable data subjectsYesChildren with SEND are explicitly vulnerable data subjects under UK GDPR Recital 75. Parents of SEND children are often in distressed or vulnerable circumstances
Innovative use of technologyYesAgentic RAG pipeline with self-correcting LLM, pgvector semantic search, automated web crawling of 152 LA Local Offers, LLM-based memory extraction from conversations
Data transfers across bordersPartialAzure OpenAI processes data in UK South region; Microsoft's DPA governs sub-processing. No data transferred outside UK/EEA by design
Processing that prevents data subjects exercising a right or using a serviceNoUsers can access the service without providing child data, though functionality is reduced

Conclusion: A DPIA is mandatory. Six of the ICO's screening criteria are met, including the processing of children's data with AI, which the ICO has explicitly stated requires a DPIA.

Data categoryLegal basisGDPR Article
User account data (email, name, password hash)Consent (registration)Art. 6(1)(a)
Child profile data (name, DOB, SEND diagnosis, EHCP status, school, key needs, provision)Explicit consentArt. 6(1)(a), Art. 9(2)(a)
Conversation contentLegitimate interest (service delivery) + ConsentArt. 6(1)(a) and Art. 6(1)(f)
AI-generated responses and RAG metadataLegitimate interest (service delivery)Art. 6(1)(f)
Knowledge base content (crawled public documents)Legitimate interest (public task information)Art. 6(1)(f)
GIAS school dataPublic task (Open Government Licence)Art. 6(1)(e)
Audit logs (IP address, user agent, actions)Legal obligation (accountability) + Legitimate interestArt. 6(1)(c), Art. 6(1)(f)
Support ticketsContract performance (service support)Art. 6(1)(b)
Safeguarding logsVital interest / Legal obligationArt. 6(1)(d), Art. 6(1)(c)

1.3 Special category data ​

The following special category data is processed under Article 9:

  • Health data: SEND diagnoses (e.g., ASD, ADHD, SEMH, SLCN), therapy details, provision information
  • Children's data: Names, dates of birth, school year, EHCP status, key needs, educational provision
  • Mental health indicators: Crisis detection system identifies users in distress or mental health crisis

Condition for processing: Explicit consent (Art. 9(2)(a)) -- collected at registration via GDPR consent flow, stored as JSONB with version, timestamp, and IP address.


Step 2: Describe the Processing ​

2.1 System overview ​

SpektraBot is an AI-powered web application that helps UK families and education professionals navigate the Special Educational Needs and Disabilities (SEND) system. It provides real-time, context-aware guidance grounded in the SEND Code of Practice, local authority Local Offers, and national charity resources.

Target users:

  • Parents and carers of children with SEND
  • Teachers and SENCOs (Special Educational Needs Coordinators)
  • System administrators

Deployment model: Cloud-native on Civo Kubernetes (London, UK data centre -- LON1 region) with UK data residency.

2.2 Data subjects ​

Data subjectRelationshipData collected
Parents/carersDirect usersEmail, name, password hash, role, persona, GDPR consent, IP address, user agent, conversation content, accessibility preferences, notification preferences
Children with SENDIndirect subjects (data provided by parents)Name, date of birth, SEND diagnosis, school name/URN, year group, EHCP status, key needs (e.g., ASD, SEMH, SLCN), provision details (e.g., speech therapy, 1:1 TA), SENCO name, key worker name, notes
Teachers/SENCOsDirect usersEmail, name, password hash, role, school name/URN, conversation content
School staff named in child profilesIndirect subjectsSENCO name, key worker name (as entered by parents)
AdvocatesDirect usersEmail, name, professional profile, case notes, referral communications

2.3 Data flows ​

2.3.1 High-level architecture ​

2.3.2 Conversation data flow (user to AI response) ​

Key privacy design decisions in this flow:

  1. Child PII is NOT sent to Azure OpenAI -- The system prompt and RAG context contain only the user's query text and retrieved public knowledge base documents. Child names, DOBs, diagnoses, and school details are NOT included in the LLM prompt.
  2. Conversation content IS sent to Azure OpenAI -- The user's message text (which may contain PII they choose to include) and the last 5 conversation messages are sent as context.
  3. Azure OpenAI processes in UK South -- Microsoft's Azure OpenAI service processes data within the UK South Azure region. Microsoft's Data Protection Addendum applies.

Confirmed: Microsoft's Azure OpenAI Service terms explicitly state that customer data (prompts and completions) is not used for model training. Azure OpenAI does not retain prompt/completion data beyond the API request lifecycle. Microsoft's Data Protection Addendum (DPA) governs all processing. Abuse monitoring is enabled by default but processes data ephemerally and does not train models.

2.3.3 Child profile data flow ​

2.3.4 Knowledge base and crawler data flow ​

2.3.5 Data sync flows (public data) ​

The following data is synced from government open data sources (Open Government Licence v3.0). This data does not contain personal data:

Data sourceFrequencyDataStorage
DfE GIAS APIDaily (3 AM UTC)52,000+ English schools with SEND provision dataschools table
Ofsted inspection dataDailyInspection ratings, framework versionsofsted_ratings table
SEN2/EHCP statisticsDailyLA-level EHCP assessment statisticsla_ehcp_statistics table
SEN primary need dataDailyLA-level SEN need breakdownsla_sen_need_breakdowns table
Tribunal statisticsDailyLA-level appeal outcomesla_tribunal_statistics table
Area SEND inspectionsDailyJoint Ofsted/CQC inspection resultsla_send_inspections table

2.4 Data storage and retention ​

Note: Retention periods below implement the Tiered Data Retention Policy, validated against ICO Children's Code Standards 7 & 8.

Data categoryStorage locationRetention periodRetention tierDeletion method
Conversations and messagesPostgreSQL (Civo UK)90 days from last activity (updatedAt)Tier 1Automated CronJob; messages cascade-delete via FK
AI response metadata (RAG traces)PostgreSQL (Civo UK)90 days from last activity (with parent conversation)Tier 1Cascade deletion with conversation
Child memoriesPostgreSQL (Civo UK)Account lifetime (closure or 12-month dormancy)Tier 2Soft delete β†’ 30-day grace β†’ hard delete. sourceConversationId set to NULL when conversation deleted (memories survive)
User accountsPostgreSQL (Civo UK)Account lifetime (user request or 12-month dormancy)Tier 3Cascade deletion of all associated data
Child profilesPostgreSQL (Civo UK)Account lifetime (with parent account)Tier 3Soft delete β†’ 30-day grace β†’ hard delete; cascade with parent
Safeguarding logsPostgreSQL (Civo UK)Indefinite (statutory: Children Act 2004 s.11, KCSIE 2024)Tier 4Manual review only β€” excluded from automated deletion
Knowledge base (crawled/uploaded)PostgreSQL (Civo UK)Until manually removed by adminN/A (no PII)Admin deletion
Audit logsPostgreSQL (Civo UK)2 years from creationN/AAutomated CronJob; denormalised email preserved so logs remain interpretable after user deletion
Support ticketsPostgreSQL (Civo UK)90 days after resolutionTier 1Automated CronJob; resolvedAt + 90 days
Database backupsCivo Object Storage (LON1, encrypted)30 days rolling (S3 lifecycle)N/AAutomatic expiry
Redis cache (sessions, WebSocket state)Redis (Civo UK, in-memory)Ephemeral (cleared on restart)N/AAutomatic eviction

2.5 Data processors and sub-processors ​

ProcessorPurposeData sharedLocationDPA in place?
Microsoft (Azure OpenAI)LLM inference (GPT-4.1) and embedding generationUser message text, conversation history (last 5 messages), knowledge base document chunks for embeddingUK South (London)Yes β€” Microsoft Online Services DPA applies automatically to all Azure services. No separate signing required.
Civo LtdKubernetes hosting, block storage, object storageAll application data (encrypted at rest on Civo infrastructure)LON1 (London, UK)Yes β€” Civo standard DPA, signed as part of service agreement. Available at civo.com/legal.
ResendTransactional email deliveryUser email addresses, email content (verification, password reset, notifications, digest)US (AWS infrastructure) β€” email addresses and content transit US serversYes β€” Resend DPA signed. UK-US Data Bridge adequacy applies. Data limited to email addresses and transactional content (no child PII).
GitHub (Microsoft)Container image registry (GHCR), source code hostingNo personal data (code and Docker images only)USN/A (no personal data)
Let's EncryptTLS certificate issuanceDomain name onlyUSN/A (no personal data)

DPA status: Microsoft DPA applies automatically via Online Services terms. Civo DPA signed with service agreement. Resend DPA signed separately. All DPAs filed in the organisation's compliance records.

2.6 Technology stack summary ​

ComponentTechnologyPrivacy relevance
FrontendNext.js 15, React 19Runs in user's browser; no server-side personal data processing
BackendNode.js 22, Express, TypeScriptProcesses all personal data; JWT authentication; rate limiting
DatabasePostgreSQL 16 + pgvectorStores all personal data; encrypted at rest; user isolation enforced
CacheRedis 7.4Ephemeral session data and WebSocket coordination only
AI/LLMAzure OpenAI GPT-4.1 (UK South)Processes conversation text; does NOT store data (per Azure terms)
EmbeddingsAzure OpenAI text-embedding-3-largeConverts text to vectors; no PII in resulting vectors
MonitoringPrometheus + GrafanaOperational metrics only; no personal data
EmailResend + React EmailSends transactional emails containing user email addresses
TLSLet's Encrypt (cert-manager)All data in transit encrypted via HTTPS/WSS
Backuppg_dump to Civo S3Contains full database snapshot including personal data

Step 3: Consultation Process ​

3.1 Internal consultation ​

ConsulteeRoleDateInput provided
[SIGN-OFF REQUIRED]Data Protection Officer[SIGN-OFF REQUIRED]Review of DPIA, advice on processing risks and mitigations
[SIGN-OFF REQUIRED]Technical Lead / CTO[SIGN-OFF REQUIRED]System architecture, security controls, data flows
[SIGN-OFF REQUIRED]Product Owner[SIGN-OFF REQUIRED]Purpose and necessity of processing, user needs
[SIGN-OFF REQUIRED]SEND Domain Expert[SIGN-OFF REQUIRED]Vulnerability of data subjects, trauma-informed design requirements

3.2 External consultation ​

ConsulteeOrganisationDateInput provided
[SIGN-OFF REQUIRED]ICO (if prior consultation required)[SIGN-OFF REQUIRED]Prior consultation under Art. 36 not expected to be required β€” residual risks are at acceptable levels. If ICO consultation is triggered, record outcome here.
[SIGN-OFF REQUIRED]Pilot school DPO(s)[SIGN-OFF REQUIRED]School-specific data protection requirements, data sharing arrangement review
[SIGN-OFF REQUIRED]Parent representatives[SIGN-OFF REQUIRED]Views on data processing, concerns, preferences (via pre-pilot consultation sessions)
[SIGN-OFF REQUIRED]SENDIASS representative[SIGN-OFF REQUIRED]Views on appropriateness of AI for SEND guidance, safeguarding considerations

3.3 Data subject views ​

Data subject views are obtained through the following channels:

  1. Pre-pilot consultation β€” Parents at pilot schools are invited to an information session explaining what data is collected, how AI processes it, and what controls they have. A feedback form captures concerns and preferences before onboarding begins.
  2. Registration consent flow β€” Granular consent options with plain-English explanations of each data category. Parents can opt out of specific data types (e.g., child memories) while retaining core functionality.
  3. In-app feedback β€” A persistent feedback mechanism allows parents to report concerns about data handling, AI responses, or privacy at any time.
  4. School DPO engagement β€” Each pilot school's DPO receives the DPIA, privacy notice, and data sharing arrangement for review before the school is onboarded. DPO feedback is incorporated into system controls.
  5. Pilot review cycle β€” At the end of each school term during the pilot, parent feedback is collected via a structured survey including data protection questions. Material concerns are addressed in DPIA updates.

[SIGN-OFF REQUIRED: Record dates and outcomes of consultation sessions as they occur.]


Step 4: Assess Necessity and Proportionality ​

4.1 Necessity of each data element ​

Data elementPurposeNecessary?Could purpose be achieved with less data?
Parent emailAccount creation, authentication, notificationsYesNo -- required for account security and communication
Parent namePersonalisation, support ticketsYesCould use pseudonym, but real name aids safeguarding
Child namePersonalisation of AI responses, child profile identificationYesCould use initials or pseudonym; full name aids parent experience
Child DOBAge-appropriate guidance (SEND provisions vary by age/key stage)YesYear group alone could suffice for most purposes
SEND diagnosisTailored guidance specific to the child's needsYesCore to the service purpose
EHCP statusProcess-specific guidance (requesting, appealing, reviewing)YesCore to the service purpose
School name/URNLocation-aware guidance (LA-specific Local Offer), school-specific policiesYesLA detection from postcode could partially substitute
Key needs (ASD, SEMH, etc.)Targeted information retrievalYesOverlaps with diagnosis but provides structured filtering
Provision detailsUnderstanding current support to identify gapsPartiallyCould be captured in conversation rather than structured profile
SENCO name / Key workerContext for advice (e.g., "ask your SENCO, [name]")PartiallyCould be omitted; adds personalisation but not essential
Conversation contentAI response generation, context continuityYesCore to the service; cannot function without it
IP addressSecurity (rate limiting, abuse detection), audit trailYesRequired for security; stored in audit logs only
Accessibility preferencesWCAG compliance, inclusive designYesCould use browser defaults, but explicit preferences improve UX

4.2 Proportionality assessment ​

Is the processing proportionate to the purpose?

The purpose of SpektraBot is to help parents of children with SEND navigate an extremely complex legal and bureaucratic system. The SEND Code of Practice is 292 pages long. Many parents are in distress, have experienced repeated institutional failures, and cannot afford professional advocates (who charge GBP 100-300/hour).

The personal data collected is the minimum necessary to provide personalised, context-aware guidance. Without knowing a child's diagnosis, EHCP status, and school/LA context, the AI system would provide only generic information that parents can already find (and struggle to interpret) on gov.uk.

Data minimisation measures in place:

  1. Child PII is NOT sent to the LLM -- The Azure OpenAI API receives only the user's message text and retrieved public knowledge base documents. Child names, DOBs, and diagnoses are kept within the PostgreSQL database.
  2. Tiered retention -- Conversation transcripts are deleted 90 days after last activity (updatedAt). Child support memories persist for account lifetime. Safeguarding records are retained indefinitely. See Data Retention Policy.
  3. Soft delete with cascade -- When a user requests deletion, all associated data (children, conversations, messages, memories, documents) is cascade-deleted.
  4. Conversation history limited -- Only the last 5 messages are sent as context to the LLM, not the full history.
  5. Privacy-preserving feedback -- The feedback_events table uses a one-way session hash instead of a user ID, making feedback anonymous by design.
  6. Accessibility preferences are non-identifying -- Font size, theme, and display preferences do not constitute personal data.

4.3 Lawfulness, fairness, and transparency ​

PrincipleHow it is met
LawfulnessProcessing is based on explicit consent (Art. 6(1)(a) and Art. 9(2)(a)). Consent is collected at registration with clear, granular options. Users can withdraw consent at any time.
FairnessUsers are informed about AI processing before they start chatting. Trauma-informed design ensures vulnerable users are not exploited or misled. Crisis detection routes users to human support services.
TransparencyPrivacy policy available at /policies/privacy. DPIA published. AI responses include source citations so users can verify information. The system does not pretend to be human.
Purpose limitationData is collected for SEND guidance only. Not used for marketing, profiling for commercial purposes, or shared with third parties beyond sub-processors.
Data minimisationSee Section 4.2 above. Each data element is justified against the service purpose.
AccuracyUsers can edit their child profiles at any time. Knowledge base content is validated by administrators. AI responses include hallucination checking.
Storage limitationTiered retention policy: conversations 90 days from last activity, memories for account lifetime, safeguarding indefinite. Backups expire after 30 days. See Data Retention Policy.
Integrity and confidentialitySee Section 6 (security measures). TLS in transit, encryption at rest, user isolation, role-based access control.

Step 5: Identify and Assess Risks ​

5.1 Risk assessment methodology ​

Risks are assessed using the ICO's recommended framework:

  • Likelihood: Remote / Possible / Probable
  • Severity: Minimal / Significant / Severe
  • Overall risk: Low / Medium / High / Very High

5.2 Risks to children (indirect data subjects) ​

#RiskLikelihoodSeverityOverallMitigation (see Step 6)
C1Child's SEND data is accessed by unauthorised person (data breach)PossibleSevereHighM1, M2, M3, M4, M12
C2AI provides incorrect guidance leading to inappropriate educational provision decisionsPossibleSignificantMediumM5, M6, M7
C3Child's SEND diagnosis or EHCP status is inferred or exposed through AI responses shared outside the platformRemoteSignificantMediumM8, M9
C4Child's data is retained beyond the necessary periodRemoteSignificantLowM10, M11
C5AI-extracted child memories contain inaccurate information that persists and affects future guidancePossibleSignificantMediumM13, M14
C6Safeguarding concern about a child is not appropriately escalatedRemoteSevereHighM15, M16

5.3 Risks to parents/carers (direct data subjects) ​

#RiskLikelihoodSeverityOverallMitigation (see Step 6)
P1Parent's conversation content (potentially distressing personal details) is accessed by unauthorised personPossibleSevereHighM1, M2, M3
P2AI provides legally inaccurate advice leading parent to take harmful action (e.g., wrong tribunal procedure)PossibleSignificantMediumM5, M6, M7, M17
P3Parent in mental health crisis does not receive appropriate signposting to emergency servicesRemoteSevereHighM15, M16, M18
P4Parent's data is used by Azure OpenAI for model training without consentRemoteSignificantMediumM19, M20
P5Parent cannot effectively exercise right to erasure due to data existing across multiple systems (DB, backups, Redis)RemoteSignificantLowM10, M11, M21
P6Parent is re-traumatised by insensitive AI responsesPossibleSignificantMediumM22, M23
P7Admin user accesses parent conversations without legitimate purposePossibleSignificantMediumM24, M25

5.4 Risks to schools and education professionals ​

#RiskLikelihoodSeverityOverallMitigation (see Step 6)
S1School staff member named in child profile has their data processed without their knowledgeProbableMinimalLowM26
S2Teacher/SENCO's conversation content about a pupil is accessed by a parent (cross-user data leak)RemoteSignificantMediumM1, M2, M3
S3AI provides incorrect information about a school's SEND provision based on outdated crawled dataPossibleMinimalLowM27, M28

5.5 AI-specific risks ​

#RiskLikelihoodSeverityOverallMitigation (see Step 6)
A1Prompt injection via crawled content causes AI to generate harmful or misleading responsesPossibleSignificantHighM29, M30
A2AI hallucination produces fabricated legal references or procedures that parent acts uponPossibleSignificantHighM5, M6, M7
A3Bias in the AI model leads to systematically different quality of advice for certain SEND categoriesRemoteSignificantMediumM31, M32
A4AI system develops over-reliance, discouraging parents from seeking professional human advicePossibleSignificantMediumM17, M33
A5LLM-extracted child memories create an inferred profile beyond what the parent explicitly consented toPossibleSignificantMediumM13, M14, M34

5.6 Infrastructure and security risks ​

#RiskLikelihoodSeverityOverallMitigation (see Step 6)
I1Database breach exposes all user and child dataRemoteSevereHighM1, M2, M3, M4, M35
I2Backup files containing personal data are accessed by unauthorised partyRemoteSevereMediumM36, M37
I3Redis cache leaks session data or WebSocket stateRemoteMinimalLowM38
I4Kubernetes cluster compromise provides access to all services and secretsRemoteSevereHighM39, M40, M41
I5Sub-processor (Civo, Microsoft, Resend) data breach affects SpektraBot usersRemoteSevereMediumM42, M43

Step 6: Identify Measures to Mitigate Risk ​

6.1 Security measures (technical) ​

IDMeasureRisks addressedStatus
M1User data isolation -- Every database query includes userId filter; enforced at service layer. WebSocket connections join user-specific rooms. Conversation access verified before any read/write.C1, P1, S2, I1Implemented
M2JWT authentication with role-based access control (parent, teacher_senco, admin). Token-based auth for all API and WebSocket connections.C1, P1, S2Implemented
M3TLS encryption in transit -- All connections use HTTPS/WSS via Let's Encrypt certificates. SSL redirect enforced at ingress.C1, P1, S2, I1Implemented
M4Encryption at rest -- PostgreSQL on Civo block storage (civo-volume) provides volume-level encryption. Backups encrypted in Civo Object Storage.C1, I1Implemented
M12Security headers -- Helmet middleware applies security headers (CSP, HSTS, X-Frame-Options, etc.). CORS restricted to specific origins.C1Implemented
M35Network security -- Kubernetes NetworkPolicy restricts pod-to-pod communication: ingress-nginx β†’ web/api only; api β†’ postgres/redis only; postgres ← backup CronJob only; deny all other intra-cluster traffic. Egress limited to Azure OpenAI, GIAS API, Resend, and DNS. Pods run as non-root (UID 1000) with dropped capabilities. No privilege escalation.I1Implemented
M36Backup encryption -- Database backups stored in Civo Object Storage with S3-compatible encryption. Backup integrity verified (rejects files < 1MB, verified with pg_restore --list).I2Implemented
M37Backup access control -- S3 credentials stored as Kubernetes secrets. Bucket access restricted to service account.I2Implemented
M38Redis security -- Redis used for ephemeral cache only (session state, WebSocket pub/sub). No personal data persisted in Redis. Redis accessible only within cluster network.I3Implemented
M39Kubernetes security context -- All pods run as non-root, all capabilities dropped, no privilege escalation allowed.I4Implemented
M40Secrets management -- All credentials stored as Kubernetes secrets (not in Helm values or source code). External secrets management.I4Implemented
M41Image security -- Private GHCR registry with imagePullSecrets. Commit-SHA image tags (mutable tags like 'latest' are blocked by Helm guard).I4Implemented

6.2 AI and data quality measures ​

IDMeasureRisks addressedStatus
M5Self-correcting RAG pipeline -- LangGraph state machine with hallucination checking and answer quality grading. Responses verified as grounded in retrieved documents.C2, P2, A2Implemented
M6Source citations -- Every AI response includes citations to the specific knowledge base documents used, enabling users to verify information independently.C2, P2, A2Implemented
M7Maximum retry limits -- Self-correction loop limited to 3 retries to prevent infinite loops and ensure response quality.C2, P2, A2Implemented
M17Disclaimer and limitations -- AI responses include clear disclaimers that SpektraBot provides information, not legal advice. Users directed to SENDIASS, IPSEA, and professional advocates for legal matters.P2, A4Implemented
M29Knowledge base validation workflow -- All crawled and uploaded content goes through admin validation (pending/approved/rejected/flagged) before being used in RAG retrieval.A1Implemented
M30Prompt injection defence -- Multi-layer approach: (1) Crawler domain whitelisting restricts sources to known LA/charity sites. (2) Content sanitisation strips HTML, scripts, and suspicious patterns during crawl ingestion. (3) Admin validation gate β€” all crawled content requires admin approval before entering RAG retrieval. (4) System prompt security markers detect attempts to override instructions. (5) Retrieved KB chunks are presented as quoted context, not as system instructions.A1Implemented
M31Evaluation test suite -- 74-case evaluation suite testing response quality across SEND categories, including trauma-informed grading.A3Implemented
M32LA coverage monitoring -- Admin dashboard tracks knowledge base coverage across all 152 English LAs. Gaps identified and addressed through targeted crawling.A3Implemented
M33Help system and onboarding -- Welcome wizard and help drawer include links to SENDIASS, IPSEA, and Spectrum Dynamics for human support. Explicit statement that SpektraBot is an AI tool, not a replacement for professional advice.A4Implemented

6.3 Children's data protection measures ​

IDMeasureRisks addressedStatus
M8Child PII not sent to LLM -- Azure OpenAI receives only user message text and public knowledge base documents. Structured child profile data (name, DOB, diagnosis) is NOT included in the LLM prompt.C3Implemented
M9No child accounts -- Children do not have user accounts. All child data is managed by the parent/carer account holder. Children do not directly interact with the system.C3Implemented
M13Memory confidence scoring -- LLM-extracted child memories include a confidence score (0.0-1.0). Low-confidence memories can be identified and reviewed.C5, A5Implemented
M14Memory source tracking -- Each extracted memory links back to the source conversation and message, enabling audit and correction.C5, A5Implemented
M34Memory deduplication -- SHA-256 content hashing prevents duplicate memories. Superseded memories are marked inactive rather than deleted (audit trail).A5Implemented
M26Third-party names -- SENCO and key worker names entered by parents are stored in child profiles only. These names are processed under legitimate interest (Art. 6(1)(f)) β€” the purpose is to personalise AI guidance (e.g., "speak to your SENCO, Mrs Smith"). LIA conclusion: Processing is minimal (name only, no contact details), proportionate (aids parent communication), and low-risk (names are already known to the parent). No separate consent required. Names are deleted when the child profile is deleted (cascade). Staff members have no account, no direct processing, and no profiling. If a named individual requests erasure under Art. 17, their name is removed from the child profile field.S1Addressed

6.4 Data subject rights measures ​

IDMeasureRisks addressedStatus
M10Tiered automated retention -- Conversations deleted 90 days after last activity (updatedAt) via daily CronJob. Child memories persist for account lifetime (soft-deleted memories hard-deleted after 30-day grace). Safeguarding records retained indefinitely. Configurable via DATA_RETENTION_DAYS environment variable. See Data Retention Policy.C4, P5Implemented
M11Right to erasure -- Users can request account deletion. Cascade deletion removes all associated data (children, conversations, messages, memories, documents, feedback). Soft delete with deletedAt timestamp.C4, P5Implemented
M21Backup erasure procedure -- Database backups are stored in Civo Object Storage with a 30-day S3 lifecycle policy. When a user exercises their right to erasure, data is deleted immediately from the live database. Backup copies naturally expire within 30 days. The ICO accepts that granular erasure from encrypted backup archives is technically disproportionate where backups have a defined, short retention period and are not used for any purpose other than disaster recovery. If a backup is restored, the data retention CronJob will re-delete any data that was erased from the live database (since the erasure itself is recorded). Privacy notice informs users that backup copies may persist for up to 30 days.P5Implemented
M24Audit logging -- 30+ action types tracked in audit logs including admin access to user data. All admin actions logged with user ID, action type, resource, IP address, and user agent.P7Implemented
M25Role-based access control -- Admin routes gated by role middleware. Non-admin users cannot access other users' data.P7Implemented

6.5 Crisis and safeguarding measures ​

IDMeasureRisks addressedStatus
M15Crisis detection system -- Automated detection of crisis indicators in user messages. Crisis module overrides normal AI response with signposting to emergency services (Samaritans, Childline, 999).C6, P3Implemented
M16Safeguarding logging -- Safeguarding concerns are logged in safeguarding_logs table with dedicated audit trail.C6, P3Implemented
M18Distress detection -- Separate from crisis detection, the system detects emotional distress and adjusts AI responses to be more empathetic, slower-paced, and focused on immediate needs.P3Implemented
M22Trauma-informed design -- All AI responses follow trauma-informed practice: validate emotions first, never blame or lecture, empower rather than prescribe, normalise the experience, respect autonomy.P6Implemented
M23Banned language patterns -- System prompt enforces a strict set of prohibited phrases ("you should", "unfortunately", "obviously", "just do this") that could re-traumatise vulnerable users. Automated grading checks for compliance.P6Implemented

6.6 Sub-processor and transfer measures ​

IDMeasureRisks addressedStatus
M19Azure OpenAI data handling -- Microsoft's Azure OpenAI service does not use customer data for model training (per Azure OpenAI terms). Data processed in UK South region.P4Implemented (by contract)
M20Data Processing Agreements -- DPAs in place with all sub-processors processing personal data. Microsoft: Online Services DPA (automatic). Civo: standard DPA (signed with service agreement). Resend: DPA signed separately. DPAs filed in compliance records and reviewed at each DPIA review.P4, I5Implemented
M42Sub-processor monitoring -- Annual review of sub-processor compliance, aligned with DPIA review schedule (Appendix E). Review covers: DPA currency, security certifications (SOC 2, ISO 27001), data processing locations, breach history, and terms changes. Any material change triggers an interim DPIA review.I5Implemented
M43Incident response and breach notification -- Breach notification procedure: (1) Contain and assess within 4 hours of detection. (2) Record breach in internal register regardless of severity. (3) If risk to data subjects: notify ICO within 72 hours via the ICO's online breach reporting tool. (4) If high risk to data subjects: notify affected individuals without undue delay, in plain English, explaining what data was affected, what we are doing, and what they should do. (5) For children's data breaches: also notify the parent/carer account holder and the relevant pilot school DPO. (6) Post-incident review within 14 days; update DPIA if systemic. PagerDuty alerts configured for automated breach detection (API/DB/backup failures).I5Implemented

6.7 Organisational measures ​

IDMeasureRisks addressedStatus
M27Knowledge base currency -- Automated crawler re-crawls LA Local Offers on admin-triggered schedule. Content hashing detects changes. Old content can be removed.S3Implemented
M28Admin validation workflow -- All content must be approved by an administrator before being used in RAG retrieval. Rejected content is excluded.S3, A1Implemented

Step 7: Sign Off and Record Outcomes ​

7.1 Summary of residual risks ​

After applying the mitigations in Step 6, the following residual risks remain:

Risk IDDescriptionResidual risk levelAccepted?Rationale
C1Child data breachMedium (reduced from High)AcceptedUser isolation, encryption at rest, TLS in transit, JWT auth, RBAC, and audit logging reduce likelihood to Possible/Remote. Residual risk is proportionate to the service benefit. Reviewed annually.
C6Safeguarding concern not escalatedMedium (reduced from High)Accepted with conditionCrisis detection is automated but not infallible. Condition: SpektraBot is not a substitute for school safeguarding procedures. Privacy notice and onboarding make clear that users should contact emergency services directly for immediate safety concerns.
P1Conversation data breachMedium (reduced from High)AcceptedSame security controls as C1. 90-day conversation retention limits the exposure window.
P3Crisis user not signpostedLow (reduced from High)AcceptedDual-layer detection (crisis + distress) with hardcoded signposting to Samaritans, Childline, 999. System is a supplement to, not replacement for, human crisis support.
A1Prompt injection via crawled contentMedium (remains High)Accepted with conditionAdmin validation of all crawled content before RAG inclusion. Domain whitelisting restricts crawl targets to known LA/charity sites. Condition: Content sanitisation (HTML stripping, script removal, chunk size limits) applied before embedding and before LLM context injection.
A2AI hallucinationMedium (reduced from High)AcceptedSelf-correcting RAG with hallucination grading, source citations for independent verification, clear disclaimers that output is informational not legal advice. Residual risk inherent in all LLM systems; mitigations reduce practical impact.
I1Database breachLow (reduced from High)AcceptedEncryption at rest, TLS in transit, Kubernetes network isolation, non-root pods, secrets management, audit logging. No public database endpoint.
I4Kubernetes cluster compromiseLow (reduced from High)AcceptedNon-root containers, all capabilities dropped, no privilege escalation, GHCR private registry with imagePullSecrets, commit-SHA image tags, RBAC. Civo manages control plane security.

7.2 Actions required before processing can commence ​

#ActionOwnerDeadlineStatus
1Sign DPA with Microsoft (Azure OpenAI)Technical LeadBefore pilot launchComplete β€” Microsoft Online Services DPA applies automatically to all Azure services
2Sign DPA with Civo LtdTechnical LeadBefore pilot launchComplete β€” Civo standard DPA signed with service agreement
3Sign DPA with ResendTechnical LeadBefore pilot launchComplete β€” Resend DPA signed separately
4Complete ICO registration[SIGN-OFF REQUIRED]Before pilot launch[SIGN-OFF REQUIRED]
5Define and implement audit log retention periodTechnical LeadBefore pilot launchComplete β€” 2-year retention, automated CronJob deletion. Denormalised email preserved in log entries so records remain interpretable after user account deletion.
6Define safeguarding log retention periodIndefinite β€” statutory requirement (Children Act 2004 s.11, KCSIE 2024). Excluded from automated deletion.Defined in Data Retention PolicyComplete
7Document breach notification procedure (72-hour ICO notification)Technical LeadBefore pilot launchComplete β€” documented in M43. Procedure: contain (4h), register, notify ICO (72h), notify affected individuals (without undue delay), school DPO notification for children's data, 14-day post-incident review.
8Implement content sanitisation for crawled KB content before LLM injectionTechnical LeadBefore pilot launchImplemented β€” crawler strips HTML tags, scripts, and style blocks during text extraction. Chunks are plain text only. Domain whitelisting restricts crawl targets to known LA/charity sites. Admin validation required before content enters RAG retrieval.
9Enable Kubernetes NetworkPolicyTechnical LeadBefore pilot launchImplemented β€” NetworkPolicy rules: ingress-nginx β†’ web (3000) and api (3001); api β†’ postgres (5432) and redis (6379); postgres ← backup CronJob; deny all other pod-to-pod traffic. Egress restricted to Azure OpenAI, GIAS API, Resend, and DNS.
10Conduct legitimate interest assessment for named school staff in child profilesTechnical LeadBefore pilot launchComplete β€” LIA conclusion in M26: processing is minimal (name only), proportionate (aids parent-school communication), and low-risk. No separate consent required. Erasure available on request (Art. 17).
11Obtain and document data subject views (parent consultation)Product OwnerBefore pilot launchProcess defined in Section 3.3. Pre-pilot information sessions, registration consent flow, in-app feedback, termly surveys. [SIGN-OFF REQUIRED: Record consultation dates and outcomes.]
12Consult pilot school DPOsProduct OwnerBefore each school onboardingProcess defined β€” each school DPO receives DPIA, privacy notice, and data sharing arrangement for review before onboarding. [SIGN-OFF REQUIRED: Record consultation dates and DPO responses.]
13Publish privacy notice referencing this DPIATechnical LeadBefore pilot launchComplete β€” privacy notice published at /policies/privacy. References tiered retention policy and DPIA. Updated to describe all data categories, retention tiers, and data subject rights.
14Document backup erasure procedureTechnical LeadBefore pilot launchComplete β€” documented in M21. 30-day S3 lifecycle policy provides natural expiry. If backup is restored, data retention CronJob re-deletes previously erased data. Privacy notice informs users of up to 30-day backup persistence.
15Establish sub-processor review scheduleTechnical LeadBefore pilot launchComplete β€” annual review aligned with DPIA review schedule (Appendix E). Covers DPA currency, security certifications, data processing locations, breach history, and terms changes.

7.3 Decision ​

Recommendation: Processing should proceed with conditions. All technical and organisational mitigations identified in Step 6 are implemented or in progress. Residual risks are at acceptable levels (Low–Medium) and proportionate to the significant benefit the service provides to vulnerable families navigating the SEND system.

DecisionProceed with conditions
Conditions(1) ICO registration completed before pilot launch. (2) Pilot school DPO consultations completed before each school is onboarded. (3) Parent consultation sessions held before onboarding families. (4) DPIA reviewed annually and after any material change to processing.
Date[SIGN-OFF REQUIRED]
Decision maker[SIGN-OFF REQUIRED]
Role[SIGN-OFF REQUIRED]

7.4 DPO advice ​

The DPO should review this completed DPIA and provide formal advice on whether processing should proceed.

DPO name[SIGN-OFF REQUIRED]
DPO advice[SIGN-OFF REQUIRED]
Date[SIGN-OFF REQUIRED]

7.5 Sign-off ​

RoleNameSignatureDate
Data Controller representative[SIGN-OFF REQUIRED][SIGN-OFF REQUIRED][SIGN-OFF REQUIRED]
Data Protection Officer[SIGN-OFF REQUIRED][SIGN-OFF REQUIRED][SIGN-OFF REQUIRED]
Technical Lead[SIGN-OFF REQUIRED][SIGN-OFF REQUIRED][SIGN-OFF REQUIRED]
Product Owner[SIGN-OFF REQUIRED][SIGN-OFF REQUIRED][SIGN-OFF REQUIRED]

Appendix A: Automated Decision-Making Assessment ​

UK GDPR Article 22 restricts solely automated decision-making that produces legal effects or similarly significantly affects individuals.

Assessment:

SpektraBot's AI generates informational guidance, not decisions. The system:

  • Does NOT determine a child's EHCP eligibility
  • Does NOT approve or reject educational provision
  • Does NOT make referrals to local authority services on behalf of users
  • Does NOT replace professional legal or educational advice

However, the guidance provided may significantly influence parents' decisions about:

  • Whether to request an EHCP assessment
  • Whether to appeal a decision to the SEND Tribunal
  • What provision to request in an EHCP annual review
  • Whether to make a formal complaint about a school

Conclusion: SpektraBot does not fall within Article 22's scope as it does not make automated decisions with legal effects. However, given the potential influence of its guidance on consequential decisions, the following safeguards are maintained:

  1. Clear disclaimers that outputs are informational, not legal advice
  2. Source citations enabling independent verification
  3. Signposting to SENDIASS, IPSEA, and professional advocates
  4. Human-in-the-loop for all consequential actions (parent must act on guidance independently)
  5. Self-correcting RAG pipeline with hallucination checking

Appendix B: International Data Transfer Assessment ​

Transfer Impact Assessment (TIA) ​

#QuestionAnswer
1Is personal data transferred outside the UK?No, by design. All primary data storage is on Civo Kubernetes in London (LON1).
2Is personal data processed outside the UK?Partially. Azure OpenAI is configured to use UK South (London) region. Conversation text is processed by Microsoft's UK South infrastructure. Microsoft's Azure OpenAI service documentation confirms that data is processed in the selected region and is not moved to other regions for inference. Microsoft may sub-process for abuse monitoring, but this is governed by the DPA and does not involve data storage outside the processing region. Resend (email delivery) processes email addresses via US-based infrastructure, covered by the UK-US Data Bridge adequacy framework.
3What personal data is transferred?User message text and conversation history (last 5 messages) are sent to Azure OpenAI for LLM inference. Knowledge base document text is sent for embedding generation.
4Is there an adequacy decision?N/A -- processing is within the UK. If Microsoft sub-processes to EEA, the UK Extension to the EU-US Data Privacy Framework may apply.
5What transfer mechanism is used?Standard Contractual Clauses (SCCs) are included in Microsoft's DPA as a fallback mechanism.
6Are supplementary measures needed?Not for UK South processing. If any sub-processing occurs outside the UK, supplementary measures should be assessed.

Azure OpenAI data handling ​

AspectDetail
Processing locationUK South (London) -- spektrabot-openai.cognitiveservices.azure.com
Model trainingMicrosoft states customer data is NOT used for training Azure OpenAI models
Data retention by MicrosoftAzure OpenAI does not store prompt/completion data beyond the API request lifecycle (per Azure OpenAI terms)
Abuse monitoringMicrosoft may process data for abuse monitoring; can be opted out for approved use cases
EncryptionData encrypted in transit (TLS 1.2+) and at rest within Azure

Abuse monitoring: Azure OpenAI abuse monitoring is enabled by default. Microsoft processes prompts and completions ephemerally for abuse detection but does not store them beyond the request lifecycle. For approved Azure OpenAI customers, abuse monitoring can be disabled via a request form β€” this is not currently applied as the default monitoring terms are acceptable (no storage, no training). Microsoft's DPA (Online Services Data Protection Addendum) governs all processing, including abuse monitoring. The DPA includes Standard Contractual Clauses as a fallback transfer mechanism.


Appendix C: Children's Data - Additional Considerations ​

ICO Children's Code (Age Appropriate Design Code) Compliance ​

Although SpektraBot is designed for parents and education professionals (not direct use by children), the following considerations apply because the system processes children's data:

StandardApplicabilityHow met
Best interests of the childYes -- all processing should be in the child's best interestSystem purpose is to help parents secure appropriate SEND provision for their child
Data protection impact assessmentYesThis document
Age-appropriate applicationPartial -- children do not use the system directly, but their data is processedSystem designed for adult users only; no child accounts; no gamification or attention-capture design
TransparencyYes -- parents should understand what happens to their child's dataPrivacy notice explains child data processing; GDPR consent collected at registration
Detrimental use of dataYes -- child data must not be used in ways detrimental to the childData used solely to provide SEND guidance beneficial to the child
Data minimisationYesOnly data necessary for personalised SEND guidance is collected (see Step 4)
Data sharingYes -- limits on sharing child dataChild data not shared with third parties. Child PII not sent to LLM.
GeolocationN/ANo geolocation data collected
Parental controlsYes -- parents control their child's dataParents create, edit, and delete child profiles. No autonomous child data collection.
ProfilingYes -- limits on profiling childrenChild memories extracted by AI constitute a form of profiling. Mitigated by source tracking, confidence scoring, and parent ability to view/delete memories.
Nudge techniquesN/ASystem does not use nudge techniques on children
Connected toys and devicesN/ANot applicable
Online toolsYesAccessibility features (WCAG 2.2, configurable font/contrast/motion) support parents with disabilities accessing their child's data

Parental responsibility ​

SpektraBot processes children's data based on the parent/carer's consent. The system assumes:

  1. The registering adult has parental responsibility for the child(ren) whose data they enter
  2. Where there are multiple adults with parental responsibility, the registering adult has authority to consent on behalf of the child

Parental responsibility declaration: The registration flow includes a checkbox confirming: "I have parental responsibility for the child(ren) whose data I will enter, or I have the consent of all persons with parental responsibility to do so." This is stored as part of the GDPR consent record (JSONB with version, timestamp, IP address).

Dispute scenarios: If a dispute arises (e.g., separated parents with shared custody), the policy is:

  1. Either parent with parental responsibility may create a profile for their child.
  2. If one parent requests deletion of a child profile created by another parent, the request is referred to the Data Controller for manual review β€” automated deletion is not appropriate where parental responsibility is shared.
  3. The system does not arbitrate custody disputes. If both parents create separate accounts, each account's data is isolated and neither parent can see the other's conversations or memories.

Appendix D: Glossary ​

TermDefinition
SENDSpecial Educational Needs and Disabilities
EHCPEducation, Health and Care Plan -- a legal document describing a child's special educational needs and the provision to meet them
SENCOSpecial Educational Needs Coordinator -- the teacher responsible for SEND in a school
SENDIASSSpecial Educational Needs and Disabilities Information, Advice and Support Service -- free, impartial advice service for parents
IPSEAIndependent Provider of Special Education Advice -- charity providing free legal advice on SEND
LALocal Authority -- the council responsible for education in an area
GIASGet Information About Schools -- DfE database of all schools in England
RAGRetrieval-Augmented Generation -- AI technique that retrieves relevant documents before generating a response
LLMLarge Language Model -- the AI model (GPT-4.1) that generates responses
pgvectorPostgreSQL extension for vector similarity search
ICOInformation Commissioner's Office -- UK data protection regulator
DPAData Processing Agreement -- contract between data controller and processor
DPOData Protection Officer
TLSTransport Layer Security -- encryption protocol for data in transit
JWTJSON Web Token -- authentication mechanism
RBACRole-Based Access Control

Appendix E: Review Schedule ​

This DPIA must be reviewed:

  1. At least annually from the date of sign-off
  2. When there is a significant change to the processing, including:
    • New data categories collected
    • New sub-processors engaged
    • Changes to the AI model or RAG pipeline
    • Expansion to new user groups or geographies
    • Security incidents
    • Changes to UK GDPR or ICO guidance
  3. Before each new school onboarding (abbreviated review confirming no material changes)
Review #DateReviewerChanges madeNext review due
1[SIGN-OFF REQUIRED][SIGN-OFF REQUIRED]Initial DPIA[SIGN-OFF REQUIRED: +12 months from sign-off date]

This DPIA follows the structure recommended by the ICO in its DPIA guidance (https://ico.org.uk/for-organisations/guide-to-data-protection/guide-to-the-general-data-protection-regulation-gdpr/data-protection-impact-assessments-dpias/). It should be read in conjunction with SpektraBot's Privacy Policy, Terms of Service, and Data Processing Agreements.

Confidential Β· Spectrum Dynamics CIC