Skip to content

Optimize watermark loading performance with PIL channel operations - #5

Open
aliciusschroeder wants to merge 1 commit into
hellloxiaotian:mainfrom
aliciusschroeder:perf/optimize-watermark-loading
Open

Optimize watermark loading performance with PIL channel operations#5
aliciusschroeder wants to merge 1 commit into
hellloxiaotian:mainfrom
aliciusschroeder:perf/optimize-watermark-loading

Conversation

@aliciusschroeder

Copy link
Copy Markdown

This PR optimizes how watermarks are loaded by leveraging PIL's native channel operations instead of pixel-by-pixel iteration. This operation is critical as it runs for every single image during both training (477 × 3,111 = 1,483,947 patches) and testing (324 images).

Benchmark Results

Tested on 150x108 image over 1000 iterations:

  • Original approach: 14.16ms ± 1.31ms
  • Optimized approach: 1.00ms ± 0.17ms
  • Performance improvement: 14.1x faster

Implementation

  • Replace pixel iteration with PIL's channel splitting
  • Modify alpha channel using point operation
  • Recombine channels with Image.merge()
  • Remove redundancy by refactoring all loads into one definition
  • Maintain identical output quality and behavior

Validation

  • Verified identical output across transparency range (0.3-1.0)
  • No regressions in watermark removal quality (PSNR/SSIM)
  • No noticeable change in memory consumption

Impact

Given the training set of 1.48M patches and test set of 324 images, this optimization can play a significant role in optimizing the processing time during a complete training cycle (100 epochs). At 1,483,947 total patch operations, each millisecond saved in the watermark loading function translates to approximately 24.7 minutes reduction in total processing time during a complete training cycle.

Replace pixel-by-pixel iteration with PIL channel operations for
watermark alpha channel modification. Benchmarks show 14.1x speedup
(14.16ms → 1.00ms) with identical output quality.

- Use Image.split() to separate RGBA channels
- Apply transparency via alpha channel point operation
- Maintain exact same transparency behavior
- Remove code redundancy by refactoring loading into one single definition
- No changes to memory footprint

Performance validated on 150x108 test image across 1000 iterations.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant