Open benchmark · Underwater Acoustics

UniqueShip: Large, public underwater acoustic target recognition (UATR) datasets for ships

2,460 hours of ship-radiated noise from 4,218 unique vessels, split by vessel ID so no ship appears in both training and test split. Sourced from the Ocean Networks Canada (ONC) repository.

Overview

Built to train generalizable UATR models using leakproof splits

UniqueShip pairs hydrophone recordings from seven ONC deployments in the Strait of Georgia (May 2016 – November 2023) with AIS vessel tracking data. Each 5-second sample is labeled with its vessel class and 17 AIS metadata fields. Unlike earlier ONC-based datasets, every split keeps each vessel in a single partition and groups background audio by day, so test accuracy reflects performance on ships the model has never heard.

The paper describes the dataset in more detail and includes results with and without data leakage, ablations on vessel diversity vs. audio duration, a metadata analysis, and additional baseline results.

3,437
Hours of ship & background audio
4,218
Unique vessels
11
Vessel classes
2.5M
5-second recordings

What makes it different?

Larger, more diverse, and free of data leakage that inflates other benchmarks

01

Leak-free splits

Vessel audio is grouped by MMSI and background by day instead of random splitting. On previous datasets, random splitting inflated accuracy by 10–48 points.

02

Largest open ONC dataset

The balanced benchmark subset alone has 4× the audio and 12× the vessels of DeepShip, and 70% more audio than the unbalanced Oceanship dataset.

03

Rich AIS metadata

17 fields per sample, including MMSI, distance to hydrophone, speed, course, length, beam, draught, and navigation status.

04

Ready-made splits

Choose anything from a 25-hour quick-start subset to the full 3,437-hour corpus, with five 80/10/10 folds.

05

Baselines included

MobileNetV3, ViT-B/16, and SwinV2 with STFT and Mel inputs. The best result is 66.5% accuracy (Swin + Mel).

06

Cleaner background class

8km ship-free radius ensures quieter ambient samples for the background class

Data releases

Current dataset splits

All splits are vessel-disjoint and include per-sample AIS metadata. Samples are 5-second clips at 20 kHz; full-length recordings are available through the codebase. Each folder contains several zipped folders, which all must be unzipped. The labels and metadata for the datasets are given in the CSV files, where each row provides the relative path of the audio/spectrogram and its corresponding metadata/label. Please follow the leakproof folds for best standardization and benchmarking across multiple models. Ship data is labeled if one ship is within 2km and no other ships are within 4km. Background is labeled if no ships are within 8km.

Data currently provided through Google Drive links in 10GB increments. Refer to the README in each Drive folder for more details, as the smaller splits are in non-independent zips while the bigger dataset is independent zips. Since the datasets are so large, we recommend rclone as a viable option to download all the data.

Diverse vessels tracked in a busy coastal shipping channel

12 Class - All Data

All data for the 12 classes and splits presented in paper. Current version = 1.0

656 GB (Unzipped, compressed .wv), 1400 GB (Unzipped, decompressed .wav), 646 GB (Zipped) 3437h, 4218 vessels, 12 classes Released September 2026
Aerial view of five vessel classes tracked in open water

5 Class - Balanced

Contains the main 5 classes (Tug/Tow, Tanker, Passengership, Cargo) and balances the total audio for each class such that they are equal. Current version = 1.0

92 GB (Unzipped), 64.8 GB (Zipped) 213h, 3175 vessels, 5 classes Released September 2026
Diverse vessels tracked in a busy coastal shipping channel

12 Class - 5 Hours Each

Contains all ship classes and balances the total audio such that it is 5 hours each class. Current version = 1.0

26 GB (Unzipped), 17.8 GB (Zipped) 60h, 4218 vessels, 12 classes Released September 2026

Benchmark results

Leaderboards

Coming soon

Compare published results across the UniqueShip dataset splits. Rankings, evaluation metrics, and submission guidance will be available following the dataset release.

Evaluation criteria and submission instructions are currently being developed.

Results preview
In development

Rankings coming soon

Inquiries

Need additional information or a different split?

Current dataset splits are available to download directly — no request or approval is required. Use this form if you have questions, need additional information, or would like to request a split that is not currently available.

  • Ask questions about the data or documentation
  • Request additional or specialized dataset splits

Citing this dataset

Reference the dataset paper

If you use UniqueShip, please cite the paper below.

@inproceedings{hashemi_2026_uniqueship,
  author    = {Hashemi, Connor and Stout, Trevor and Hoogs, Anthony and Parham, Jason},
  title     = {UniqueShip: Mitigating Data Leakage in Acoustic Ship Classification Benchmark Datasets},
  booktitle = {OCEANS 2026},
  year      = {2026},
  pages     = {TODO}
}

Questions, corrections, or collaboration?

Our team can help with access and research partnerships.

Contact the team