Skip to main navigation Skip to search Skip to main content

Transfer Learning with Self-Supervised Vision Transformer for Large-Scale Plant Identification

  • Mingle Xu*
  • , Sook Yoon
  • , Yongchae Jeong
  • , Jaesu Lee
  • , Dong Sun Park
  • *Corresponding author for this work
  • Jeonbuk National University
  • Mokpo National University
  • Rural Development Administration

Research output: Contribution to journalConference articlepeer-review

Abstract

This paper is a working note for the PlantCLEF2022 challenge aiming to identify plants with a large-scale dataset, several millions of images and 80,000 classes. Although there are many images, each class only includes 36 images around on average and thus it can be regarded as a few-shot image classification. To address this issue, transfer learning is validated to be useful in many scenarios and a popular strategy is employing a convolution neural network (CNN) pretrained in a supervised manner. But inspired by the literature on computer vision, we instead leverage a self-supervised vision transformer (ViT) and secure the first place with MA-MRR 0.62692, 0.019 higher than the second place, and 0.116 than the third. Furthermore, we achieve 0.64079 if training the model twenty epochs longer. Compared to the popular strategy with CNN, self-supervised ViT has two advantages. First, ViT does not embrace any inductive bias, such as translating invariance embraced in CNN, and thus owns a more powerful model capacity. Second, self-supervised pretraining obtains a task-agnostic feature extractor that may be better for the downstream task. To be more specific, a recently proposed self-supervised ViT model pretrained in ImageNet, masked autoencoder (MAE), is finetuned in PlantCLEF2022 dataset and then tested to report the evaluation. Except for the challenge, we discuss its possible impacts, such as taking the dataset to pretrain a model for plant-related tasks. Especially, our preliminary results suggest that the pretrained model in PlantCLEF2022 essentially contributes to image-based plant disease recognition on several public datasets. Via our analysis and experimental results, we believe that our work encourages the community to utilize the self-supervised ViT model, the PlantCLEF2022 dataset, and our pretrained model in the dataset. Our codes and trained model are public at https://github.com/xml94/PlantCLEF2022.

Original languageEnglish
Pages (from-to)2238-2252
Number of pages15
JournalCEUR Workshop Proceedings
Volume3180
StatePublished - 2022
Event23rd Working Notes of Conference and Labs of the Evaluation Forum, CLEF 2022 - Bologna, Italy
Duration: 2022.09.52022.09.8

Keywords

  • computer vision
  • image classification
  • plant identification
  • self-supervised
  • transfer learning
  • vision transformer

Quacquarelli Symonds(QS) Subject Topics

  • Computer Science & Information Systems

Fingerprint

Dive into the research topics of 'Transfer Learning with Self-Supervised Vision Transformer for Large-Scale Plant Identification'. Together they form a unique fingerprint.

Cite this