A 66-person maritime survey found no trust difference between two AI collision scenarios
An AI assistant on a ship's bridge does not need blind trust. It needs an officer who knows when to use it and when to ignore it.
A new survey of maritime professionals tested that relationship with two simulated collision-avoidance scenarios. The first showed a relatively clear encounter. The second added four vessels, a fishing boat and a nearby buoy. Overall trust in the assistant did not differ between them (p=0.936).
Respondents were open to maritime technology and did not describe the assistant as broadly unreliable or deceptive. They also resisted changing their own decision because of its output. The group accepted support without surrendering judgment.
The exercise tested a correct assistant
Researchers collected responses from December 2025 through mid-April 2026. Of 166 people who opened the survey, 66 completed it. About half were captains or officers of the watch. Among the 62 who reported sea experience, 31 had more than 72 months.
Participants first saw electronic chart and automatic identification system views of a traffic situation. They wrote what they noticed and what action they would consider. The survey then showed the encounter again with an explanation layer displaying the vessel parameters speed and rudder with importance scores. A captain supplied the intended interpretation under the collision rules.
Static screens simulated the assistant through a Wizard-of-Oz setup. It never offered deceptive behavior or a wrong recommendation. The survey therefore measured attitudes toward plausible, correct support. It did not test what happens after a dangerous error.
After each scenario, participants completed a trust scale. Negatively worded items about unreliability, deception and suspicion had a median of 2 on a five-point scale, so respondents generally disagreed. Positive trust items had a median of 3. That is neutral, not strong endorsement.
The trust scale had high internal consistency in both scenarios, with Cronbach's alpha of 0.910 and 0.936. The altered explanation scale was less cohesive at 0.633 and 0.609. Factor analysis found three latent factors. Explanation quality did not behave as one simple attitude.
Dense traffic did not lift trust
The denser scenario did not materially change overall trust. Sea-experience groups showed no significant difference either. Age-group differences were also nonsignificant, although the second scenario came closer at p=0.062 and had substantial variation within groups.
Technology anxiety did not differ significantly by age (p=0.708) or sea experience (p=0.219). The sample is too small and uneven for sweeping demographic conclusions. It does not support treating age as a shortcut for resistance to automation.
One explanation item says more than the aggregate trust score. Participants gave a median of 2 when asked whether they would change their decision based on the agent's output. They could understand the display without treating it as an order. In the first scenario, the reverse-coded ease-of-understanding item averaged 4.1.
The assistant can organize traffic information or confirm a judgment while the officer keeps responsibility. That is a more useful design target than maximizing a trust score.
Seafarers named concrete failure modes
Open responses credited the assistant with supporting situation awareness, confidence and time savings. Some participants thought it could help a young officer check a decision.
Their concerns were operational. Faulty sensors, GPS spoofing or jamming could corrupt the inputs. Sequential traffic situations may not fit isolated snapshots. Respondents also named distraction, alarm burden and information overload.
Skill loss sat underneath those concerns. Junior officers who follow recommendations too readily may get less practice judging traffic. A tool that improves today's routine decision can weaken the human backup required for an unusual one.
READY's human-review benchmark measures a related tradeoff. Review takes time, but removing it does not remove uncertainty. On a bridge, the reviewer is also exercising a skill that can decay.
Responsibility remains with people
The International Maritime Organization defines a Maritime Autonomous Surface Ship as one that can operate independently of human interaction to some degree. Its non-mandatory MASS Code took effect on July 1, 2026. It requires a description of operating modes, and the IMO says responsibilities for masters, remote operators and onboard crew need to match the autonomy level.
The COLREGs still govern collision avoidance, including safe speed, collision risk and steering conduct. An explanation display cannot replace that framework. It has to help a responsible operator apply it under time pressure.
The survey used correct advice and simplified displays. Production tests should include late, incomplete and conflicting evidence. The pattern used by TechforHumans' replay gate fits here: rerun important cases whenever a model, prompt, tool or interface changes.
Test the handoff
A satisfaction score will not answer the operational questions. A useful test should:
- Show the source and age of every input, including missing position, speed or heading data.
- Keep advice visually separate from authority, with the responsible person and override clear.
- Run sequences of changing encounters instead of isolated screenshots.
- Record whether operators notice errors, challenge advice, recover and retain manual skill.
- Test explanation timing, since detail that helps during analysis can distract during a close encounter.
Do not let the assistant grade its own reliability. In another study, direct model self-reports correlated at only r=0.04 with observed harmful behavior. Confidence language is not evidence that advice deserves obedience.
Where the evidence stops
The final sample was 66, down from 166 survey visitors, with an uneven age distribution. Static images may have reduced realism and contributed to dropouts. The assistant never lied, missed a vessel or received spoofed data. Trust after those failures could look different.
The survey participants accepted AI as support while keeping professional judgment. Its next test is already visible in their concerns: feed the assistant stale sensor data during a changing encounter, then measure whether the officer catches the error without losing the traffic picture.



