Software · Systems · Research
Interface sounds
Ambient music
a.
Adjie Workspace
Interface sounds
Ambient music

Scripted introduction. The project study is Adjie AI’s prepared answer; use Ask AI for your own questions.

What improved tomato detection: the architecture changes or the ensemble?

Adjie AI

The study evaluates both, separately: a YOLOv11 baseline, modified single models, then three-model Weighted Boxes Fusion. The comparison keeps architectural gains distinct from gains that require ensemble inference.

Applied computer vision / Master’s thesis / 2026

TomatoVision

One scene.
Different hypotheses.

A computer-vision study that detects three greenhouse tomato maturity stages and compares a YOLOv11 baseline with modified models and a three-model WBF ensemble.

YOLOv11 baseline detections on a photograph of a tomato plant in a garden greenhouseInspect original ↗
YOLOv11 baselineQualitative demo · not validation imagery

The red fruit is classified as orange in this scene. Domain shift is shown as observed, not retouched. Photo: Kolforn; detections overlaid. Image adaptations shared under CC BY-SA 4.0.

My contribution
Computer vision research · model development · evaluation
Built with
YOLOv11 · Swin Transformer · PyTorch · OpenCV · Weighted Boxes Fusion
The problem

Start with the work.

Greenhouse tomatoes overlap, hide behind leaves, and appear at different scales, making maturity detection difficult for a single detector configuration.

What I built

Adjie evaluated a YOLOv11 baseline, Swin Transformer and multi-scale SPPF variants, then fused their predictions with Weighted Boxes Fusion while reporting single-model and ensemble results separately.

01 / Research method

Change the architecture.
Then test the ensemble.

  1. Baseline

    YOLOv11

    A fixed reference for the architectural comparisons.

  2. Single-model variants

    Swin-T
    + multi-scale SPPF

    Evaluate the modified detector independently.

  3. Prediction fusion

    Weighted
    Boxes Fusion

    Combine model outputs and report the ensemble separately.

02 / Quantitative evaluation

The result depends
on what you run.

Single-model and ensemble results from the recorded thesis evaluation. These figures do not measure the demo photograph above.

Seven evaluated configurations
ConfigurationTypemAP@0.5mAP@0.5:0.95ms / image
YOLOv11single0.7950.47055.28
+ Swin-Tsingle0.8050.47453.46
+ Swin-T + MS-SPPFsingle0.8070.47752.77
Combine 1YOLOv11 + Swin-Tensemble0.8140.49069.88
Combine 2YOLOv11 + Swin-T+MS-SPPFensemble0.8170.49269.37
Combine 3Swin-T + Swin-T+MS-SPPFensemble0.8120.48767.17
Combine 4all threeensemble0.8240.49988.93

Timing belongs to the recorded evaluation setup; it is not a browser benchmark or a deployment guarantee.

03 / Dataset & training

Keep the denominator
in the story.

Source imagery
1,051 images · Known-You Seed Co., Pingtung, Taiwan
Source annotations
12,168 green · 14,640 orange · 13,989 red
Exported split
exported 2,193 train (augmented) / 160 val / 160 test at 1280×1280
Training configuration
1088px · batch 4 · 400 epochs · patience 100
Evidence / verification

The public record.

YOLOv11 baseline
0.795 mAP@0.5
Best modified model
0.807 mAP@0.5
Three-model WBF
0.824 mAP@0.5
WBF stricter metric
0.499 mAP@0.5:0.95
Want to go deeper?