Primary Supervisor
- Position
- Research Fellow
- Division / Faculty
- Faculty of Engineering
Other QUT supervisors
- Position
- Head of School, Electrical Engineering and Robotics
- Division / Faculty
- Faculty of Engineering
Overview
Every trademark logo is tagged with "Vienna codes" — a standard list of about 1,500 categories describing what's in the image (a star, an animal, a crown, a certain style of lettering, and so on). Today, experts assign these tags by hand, which is slow and needs training. This project explores how to tag logos automatically without teaching the system from labelled examples. The idea: an AI model looks at a logo and describes it in words, and those words are then matched to the official descriptions of each Vienna category. The main challenge is that the system is good at guessing the broad type of image but struggles to pick the exact right tag, even when that tag is among the options it considered. The project asks how much we can improve this — using smarter matching and reranking tricks.
Research engagement
In this project you'll help build a system that automatically tags trademark logos with the right categories. You'll use modern AI models that can "look" at an image and describe it in words, then work on matching those descriptions to the roughly 1,500 official logo categories. The main challenge — and the heart of your work — is ranking: making sure the correct categories come out on top of the list, not buried below wrong ones. Since each logo usually belongs to several categories at once, arranged from broad types down to fine details, you'll also work on choosing how many tags to predict and how to measure whether your choices are right. Along the way you'll run experiments on the group's shared computers, compare different ideas fairly, and use a visual tool to see which logos the system gets right or wrong. The work is mostly in
Python, building on an existing, well-documented codebase rather than starting from scratch, and you'll get regular one-on-one guidance as part of an active research group — with a chance to contribute to a published paper.
Research activities
The central aim of this project is to propose a new method that outperforms current approaches to zero-shot trademark (Vienna code) classification, and to write it up as a draft ready for publication. Concretely, we expect the project to deliver:
- A novel technique that improves on the existing pipeline — most likely a better way of ranking the correct categories to the top (for example, a new reranking or matching approach) — while keeping the system free of hand-labelled training data.
- Clear evidence that it works, through fair, well-designed experiments showing measurable gains over the current best results and sensible comparisons (ablations) that explain why the method helps.
- A publication-ready draft — a written paper with results tables, analysis, and figures — together with clean, reproducible code.
Research skills
The student will gain hands-on experience with modern vision-language models, text embeddings, and semantic similarity search, and learn how to design and evaluate multi-label, hierarchical classification systems. They will also build core research skills — running systematic experiments, analysing results, and writing up findings for academic publication.
Outcomes
The central aim of this project is to propose a new method that outperforms current approaches to zero-shot trademark (Vienna code) classification, and to write it up as a draft ready for publication. Concretely, we expect the project to deliver:
- A novel technique that improves on the existing pipeline — most likely a better way of ranking the correct categories to the top (for example, a new reranking or matching approach) — while keeping the system free of hand-labelled training data.
- Clear evidence that it works, through fair, well-designed experiments showing measurable gains over the current best results and sensible comparisons (ablations) that explain why the method helps.
- A publication-ready draft — a written paper with results tables, analysis, and figures — together with clean, reproducible code.
Skills and experience
Essential
- Solid Python programming skills and comfort working with an existing codebase.
- Foundational knowledge of machine learning / deep learning (from coursework or projects).
- Comfort with the command line and running code on remote/shared computers.
- Self-motivated and able to work independently, with good written communication for the eventual paper.
Highly desirable
- Experience with deep learning frameworks (PyTorch) and the Hugging Face ecosystem (transformers, sentence-transformers).
- Exposure to any of: vision-language models, image classification, text embeddings, information retrieval / ranking, or multi-label classification.
- Familiarity with running experiments on GPUs.
Ideal candidate
- A student genuinely keen on research and publication, motivated to push for a result strong enough to submit to a conference or journal.
- Curious about multimodal AI (models that combine images and text) and happy to read recent papers, try ideas, and iterate based on results.
- Prior interest or coursework in computer vision, natural language processing, or multimodal learning is a strong plus.
Start date
2 November, 2026End date
19 February, 2027Location
GP-S825
Keywords
Contact
Osman Tursun
0416920626
osman.tursun@qut.edu.au