Skip to main navigation Skip to search Skip to main content

PViT: pooling vision transformer for active trachoma image classification

Mulugeta Shitie Zewudie, Shengwu Xiong, Xiaohan Yu*, Xiaoyu Wu, Aminu Onimisi Abdulsalami, Mengjiao Wang

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference proceeding contributionpeer-review

Abstract

Trachoma, a preventable eye infection, poses a significant risk of blindness without timely detection and treatment. In response to this urgent need, this paper develops PViT, a novel hybrid convolutional neural network, and transformer framework that integrates an enhanced multilayer perceptron (MLP) module for trachoma image classification. To address the limitations of existing approaches in capturing critical global features, the proposed enhanced MLP module enables global information pooling and local information extraction toward compressive representation learning. Evaluations on two publicly available datasets, the active trachoma inverted eyelid image dataset, and the glaucoma OCT scan dataset, verify the effectiveness of the proposed PViT model, reaching 90% accuracy that surpasses current state-of-the-art deep learning models. Ablation studies further show the significance of the pooling technique in the enhanced MLP module.

Original languageEnglish
Title of host publicationAdvanced Data Mining and Applications
Subtitle of host publication20th International Conference, ADMA 2024, Sydney, NSW, Australia, December 3-5, 2024, proceedings, part IV
EditorsQuan Z. Sheng, Gill Dobbie, Jing Jiang, Xuyun Zhang, Wei Emma Zhang, Yannis Manolopoulos, Jia Wu, Wathiq Mansoor, Congbo Ma
Place of PublicationSingapore
PublisherSpringer, Springer Nature
Pages266-280
Number of pages15
ISBN (Electronic)9789819608409
ISBN (Print)9789819608393
DOIs
Publication statusPublished - 2025
Event20th International Conference on Advanced Data Mining Applications, ADMA 2024 - Sydney, Australia
Duration: 3 Dec 20245 Dec 2024

Publication series

NameLecture Notes in Artificial Intelligence
PublisherSpringer
Volume15390
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference20th International Conference on Advanced Data Mining Applications, ADMA 2024
Country/TerritoryAustralia
CitySydney
Period3/12/245/12/24

Keywords

  • Active Trachoma
  • Image Classification
  • Pooling
  • Vision Transformer

Fingerprint

Dive into the research topics of 'PViT: pooling vision transformer for active trachoma image classification'. Together they form a unique fingerprint.

Cite this