Ask what data a service robot collects and the first answer will usually be a list of sensors: cameras, microphones, lidar, proximity sensors, touch controls and location systems.
That is only the visible layer.
A robot also generates maps, event logs, confidence scores, interaction histories, alerts, maintenance records and records of how people move around it. Cloud services may derive speech transcripts, object labels, identity matches, sentiment indicators or predictions about behaviour. The manufacturer may receive technical telemetry, while the deploying organisation sees a separate dashboard. A third-party model provider may process prompts or images without appearing in the physical product description.
The result is a data environment, not a single data collection event.
This is the second article in our Robots in the Real World series. The series governance guide explains why each deployed function needs its own Robot Function Card. Here, we examine the personal data, inferred data and operational records that card needs to uncover.
A robot is a moving data-collection point
A fixed camera has a defined field of view. A smart speaker usually sits in one room. A service robot can move between spaces, change its orientation, approach people and combine several sensors as it acts.
That mobility changes the privacy analysis. The device may enter a staff-only area, a resident's bedroom, a changing space, a clinical zone or a quiet corner where people reasonably expect more privacy. Its field of view may include documents, screens, medication, family photographs, religious items, assistive equipment or other contextual details that reveal more than the robot's task requires.
The system may also collect differently depending on what happens around it. A collision-avoidance camera might retain nothing during normal operation but save a clip after an impact. A voice system may buffer audio locally while waiting for a wake word. A support function may activate remote viewing when a fault is reported. An emergency feature may begin recording only after a distress signal.
An inventory that says camera data or audio data does not capture these conditions. It needs to record when collection begins, what triggers retention, who receives the data and what later processing takes place.
Map six layers of data
A useful inventory separates at least six data layers.
| Layer | Examples | Why it can be missed |
|---|---|---|
| Direct inputs | Voice, images, video, touch entries, identity details, commands | Often listed by sensor rather than by person or purpose |
| Environmental observations | Room layout, objects, screens, documents, noise, routes, device identifiers | Described as mapping or navigation data even where it reveals people or habits |
| Derived and inferred data | Transcripts, emotion or urgency labels, activity patterns, age bands, risk scores | May exist only in a model output or vendor dashboard |
| Generated outputs | Responses, recommendations, alerts, routes, task assignments | Treated as system output even where tied to an identifiable person |
| Operational telemetry and logs | Battery, errors, access records, update history, location, override events, diagnostic clips | Labelled technical data although it can reveal staff actions and service events |
| Improvement and training data | Sample interactions, labelled clips, prompts, corrections, feedback, model-evaluation records | Hidden inside product-improvement, quality or support clauses |
These layers should be mapped for each function. The same video frame might be used momentarily for obstacle avoidance, retained after an incident, reviewed by a support engineer and later selected for model evaluation. Those are separate processing operations even though the data originated from one sensor.
The sensor tells you how data enters the system. It does not tell you what the system will infer, retain, share or learn from it.
Include everyone in the robot's operating space
The intended user is only one affected group.
Workers may appear in video, speak near the device, carry authentication tokens, receive tasks or become visible through response-time and override logs. A robot introduced to reduce manual effort can quietly become a source of worker-performance data.
Visitors and customers may be observed before they know the robot is active. They may not understand whether it is navigating, recording, identifying or responding. In a shop, hotel or reception area, avoiding the device may be difficult without avoiding the service.
Residents, patients and care recipients may be observed in highly private situations. Their data can reveal health, mobility, distress, relationships, routines and the support they receive, even where the product is marketed as a general assistant.
Children may interact with a robot as though it were a character or companion and disclose information readily. Height, voice and behaviour can also affect how well a system recognises or responds to them.
Bystanders may not interact at all. A family member on a video call, a delivery driver at the door, a colleague passing through a corridor or a person reflected in a screen can still enter the data flow.
The inventory should therefore be organised by affected group as well as function. This helps the organisation identify people who did not choose the interaction and may not receive the same information as the intended user.
Raw sensor data may reveal more than its label suggests
Context changes the meaning of apparently ordinary data.
A room map can reveal that someone sleeps downstairs or uses mobility equipment. Repeated movement data can show religious observance, work patterns or declining activity. Background audio can reveal health discussions or union activity. A list of objects can expose medication, alcohol use or political material. Device proximity can indicate who spends time together.
Some sensor data can support biometric processing, but the legal category needs precision. An image of a face is personal data where the person is identifiable; it becomes biometric data under the GDPR's special-category rule where specific technical processing is used for the purpose of uniquely identifying that person. Similarly, a system can generate health-related data from apparently non-medical inputs where it infers a person's physical or mental condition. Inferences that reveal religious belief, trade-union membership or political opinion are also special-category data under Article 9(1), even where their source inputs look unremarkable.
The organisation should inventory both what the supplier says it collects and what can reasonably be revealed through combination, repetition and context. This is especially important where a robot operates in homes, care settings, workplaces or locations associated with health or belief.
Inferences can be more consequential than recordings
Raw data often has a short life. The inference may persist.
A spoken request can become an urgency score. Video can become a fall alert, estimated age, attention measure or behavioural classification. Movement can become a productivity indicator. Interaction history can become a prediction about whether a person needs help.
The inventory should record:
- the inference or label produced
- the source data used
- the person or group it relates to
- the confidence threshold
- where the output is displayed or sent
- what action or decision may follow
- whether the inference is retained after the source data is deleted
- whether staff can see uncertainty, correct the result or add context
An organisation may believe it has minimised data because it does not retain video, while retaining a long history of alerts and behavioural labels generated from that video. Data minimisation must examine the whole result, not just the heaviest file.
The lawful-basis and fairness questions for these operations are addressed in Service Robots and GDPR: Lawful Basis, Consent and Individual Rights.
Separate local, edge and cloud processing
Phrases such as processed on device need to be tested.
Some functions may genuinely run locally without sending content away. Others may process locally first, then transmit selected events, embeddings, thumbnails, transcripts or diagnostic samples. A robot may use one cloud for fleet management, another for speech recognition and a third-party model for generated conversation.
For each data type, the flow map should show:
- where the data is first captured
- what processing occurs on the device or local network
- what leaves the location and under which trigger
- which vendor service and subprocessor receives it
- where the service is hosted and supported
- what returns to the robot or dashboard
- what is retained at each point
- which remote users can access live or stored data
This matters for security, transparency, international transfers, processor terms and incident response. It also matters operationally: a function that depends on a remote service may behave differently or fail entirely when connectivity is lost.
The EDPB's guidance on virtual voice assistants is useful beyond smart speakers because it examines activation, multiple users, voice data, transparency and the relationship between device access and subsequent processing. Its exact legal application remains fact-specific, particularly where national rules implementing Article 5(3) of the ePrivacy Directive are engaged.
Treat telemetry as potentially personal
Technical data is not automatically non-personal.
A device identifier linked to a resident's room, a route log tied to a named worker's shift or a maintenance record containing a conversation clip can all relate to identifiable people. Remote-access logs can reveal which employee dealt with an event. Battery and fault records can expose service patterns when combined with location and time.
Telemetry often serves several parties. The deploying organisation may need it for safety and support. The manufacturer may use it to maintain the fleet, defend claims or improve products. A reseller may operate the support desk. Each purpose needs to be visible in the inventory and role map.
The EU Data Act also creates access and use rules for data generated by connected products and related services. It does not displace the GDPR. Where connected-product data is personal, both frameworks may need to be satisfied.
Examine secondary use and model improvement
Product-improvement wording deserves close attention because it can cover very different practices. Our guide to defensible vendor privacy lifecycles explains why those purposes and controls must remain visible throughout the relationship.
Aggregated error rates may create limited privacy risk. Retaining selected interactions to train a recognition model is a different operation. Human review of audio, video or prompts can expose content to people who were never apparent at the point of collection. A global fleet dataset may combine deployments from care, retail and domestic environments.
Ask the supplier to distinguish:
- service delivery for the customer
- security and fault diagnosis
- product analytics
- model evaluation
- supervised labelling or human review
- training, fine-tuning or retrieval use
- development of unrelated products
The answer should identify whether the use is mandatory, configurable or optional; whether the data is personal; which organisation determines the purpose; how material is selected; and when it is deleted.
The EDPB's Opinion 28/2024 on AI models makes clear that model anonymity and legitimate-interest assessments are case-specific. Removing obvious identifiers from a training set does not by itself establish that the resulting model or processing is anonymous.
Do not let retention hide inside event design
Retention is often distributed across the system.
The device may overwrite its buffer after minutes. The alert platform may keep a thumbnail for 30 days. The vendor's support ticket may retain the same image for years. A model-evaluation dataset may have a separate lifecycle, while audit logs are preserved for accountability or legal claims.
The inventory should therefore pair every data object with a location, purpose, owner and deletion rule. It should include backups, exported reports, diagnostic bundles and locally downloaded files. Retained according to policy is not enough where several organisations apply different policies.
Trigger-based recording also needs examination. If footage is retained only after a collision, distress alert or manual escalation, the organisation should test false triggers and confirm what happens immediately before and after the event. Pre-event buffers can contain more personal data than the final alert suggests.
Design transparency for a moving, multi-user system
A privacy notice on a website will not reach everyone encountered by a robot.
Transparency may need several layers: signs at entrances, a visible indicator on the device, concise spoken or on-screen information, staff explanations, a detailed online notice and an accessible route for questions or objections. The layer should tell people what the robot is doing now, not merely what the product is capable of doing in some configuration.
The organisation should avoid signals that imply more certainty or intelligence than the system has. A conversational style, face-like screen or personalised response can encourage people to disclose more than they would to an ordinary device. Conversely, a friendly appearance should not obscure recording, remote access or automated analysis.
The strongest transparency control is still functional restraint. If the purpose does not need audio, identity, emotion analysis or retained video, disabling the feature is more reliable than explaining unnecessary collection at length.
Build a data inventory that can survive change
Robot data maps become stale quickly because software, integrations and operating contexts change.
The Function Card should therefore record change triggers such as:
- a new sensor or previously disabled feature being activated
- a different cloud, model or subprocessor
- expansion into a new room, site or public area
- use with children, patients or another affected group
- a new dashboard or analytics function
- retention of data that was previously processed transiently
- use of service data for training or product improvement
- linking robot data to HR, customer, health or access-control records
- a change in remote-support access
These triggers should connect to change management and vendor notices. Otherwise, privacy teams learn about the new data flow after an incident, complaint or renewal. They should also feed the organisation's DPIA and AI governance assessment process.
Practical questions for the DPO and service owner
For each function, ask:
- Which people can enter the sensor range, including those who never interact?
- What raw, inferred, generated, telemetry and training data exists?
- Which data is transient, and which output or event record survives?
- What is processed locally, remotely and by third parties?
- Can optional collection, analytics and model-improvement uses be disabled?
- How are people told what is active in the place where it happens?
- Can the organisation retrieve, correct, restrict and delete data across every system?
- What happens to data when the device is replaced, returned, resold or redeployed?
These questions provide the factual base for lawful basis, DPIA, security, rights and vendor assessments. Without it, those later conclusions rest on a product brochure rather than the deployed system.
The XpertDPO view
The privacy risk of a service robot is not measured by how human it looks. A plain mobile unit can create a detailed behavioural record; a sophisticated humanoid device may run a narrowly bounded function locally.
What matters is the combination of reach, context, inference, retention and consequence. Organisations need a living data inventory that follows information from the physical environment through the device, cloud, dashboard, support chain and any later training use.
Our AI Governance and DPIA Lifecycle Support helps organisations turn that map into defensible classifications, controls and reassessment triggers. Our DPO Support helps privacy teams challenge claims, advise independently and monitor whether the approved data use matches real operation.
Next in the series, Service Robots and GDPR: Lawful Basis, Consent and Individual Rights applies the GDPR tests to the data and functions mapped here.
Sources and further reading
- GDPR, Regulation (EU) 2016/679
- EDPB Guidelines 02/2021 on virtual voice assistants
- EDPB Opinion 28/2024 on data-protection aspects of AI models
- EDPB summary of Guidelines 3/2019 on processing through video devices
- ePrivacy Directive, Article 5(3)
- EU Data Act overview
- Irish DPC guidance on data-protection impact assessments