my-small_diff-model

my-small_diff-model is an experimental diffusion model trained as part of a hands-on study of diffusion model training, optimization, and image generation workflows.

The model was trained for 50 epochs on a single GPU, reaching 3,149 optimization steps. Training behavior, learning-rate scheduling, and periodic image generation were monitored throughout the run.

This repository is intended primarily as a reproducible machine learning engineering experiment rather than a production-ready generative model.

Model Details

Model Description

This model was trained using the Hugging Face diffusers ecosystem as an experimental small diffusion model.

The project focuses on understanding the practical behavior of diffusion training, including:

  • convergence over repeated epochs
  • learning-rate scheduling
  • training stability
  • single-GPU training
  • periodic qualitative sampling
  • packaging and publishing a diffusion model through the Hugging Face Hub

Developed by: Priyanka / pynk17

Model type: Diffusion-based generative image model

Task: Image generation

Framework: PyTorch / Hugging Face Diffusers

Training hardware: Single GPU

Training epochs: 50

Total optimization steps: 3,149

Repository: https://huggingface.co/pynk17/my-small_diff-model

License: Not specified

Base model: Not documented in the available training metadata

Intended Uses

Direct Use

The model can be used for experimentation with diffusion-model inference and for studying the behavior of a relatively small generative model trained from a limited training setup.

Possible uses include:

  • studying diffusion inference
  • experimenting with sampling behavior
  • investigating the relationship between training loss and generated output quality
  • testing Hugging Face Diffusers workflows
  • educational demonstrations of diffusion training and deployment
  • reproducibility experiments

Downstream Use

The model may also be useful as a starting point for:

  • additional fine-tuning
  • inference optimization experiments
  • mixed-precision comparisons
  • memory optimization studies
  • checkpointing experiments
  • scheduler comparisons
  • model deployment experiments

The model should be treated as an experimental research artifact rather than a production model.

Out-of-Scope Use

This model has not been validated for safety-critical, commercial, or production deployment.

It should not be assumed to provide:

  • photorealistic image generation
  • robust prompt following
  • unbiased outputs
  • production-level reliability
  • safety filtering
  • factual or semantically reliable image generation

Training Details

Training Data

The model was trained using a custom image training dataset supplied to the training pipeline.

Further dataset documentation should include:

  • dataset name
  • number of training images
  • image resolution
  • preprocessing steps
  • captioning or conditioning strategy, if applicable
  • dataset license and provenance

These details are not available in the current training metadata and should be added when confirmed.

Training Procedure

Training was launched on one GPU and executed for 50 epochs.

Each epoch contained approximately 63 training iterations, resulting in a final training step of 3,149.

The run also included periodic image-generation passes. The logs show 1,000-step generation/sampling loops occurring at several points during training, each executing at approximately 34.8 iterations per second.

Optimization Behavior

Training began with a learning rate of approximately:

1.26e-5

The learning rate increased during the early training phase, reaching approximately:

1.0e-4

around epoch 7.

It then gradually decayed through the remainder of training until reaching:

0

at the end of epoch 49.

This indicates a learning-rate schedule containing an initial warm-up phase followed by gradual decay.

Training Loss

The reported training loss decreased substantially during training.

Early training:

Epoch Step Loss Learning Rate
0 62 0.358 1.26e-5
1 125 0.125 2.52e-5
2 188 0.104 3.78e-5
3 251 0.0252 5.04e-5
4 314 0.0166 6.30e-5

During later training, losses frequently fell below 0.02, although noticeable fluctuations remained throughout the run.

Examples include:

Epoch Step Loss Learning Rate
14 944 0.00669 9.32e-5
26 1700 0.00594 5.73e-5
32 2078 0.00350 3.52e-5
39 2519 0.00656 1.33e-5
47 3023 0.00176 5.57e-7
49 3149 0.0384 0

The lowest reported loss in the supplied logs was approximately:

0.00176 at epoch 47

Loss was not monotonically decreasing. Temporary increases occurred at several points, including epochs 7, 11, 18, 19, 22, 33, and 43.

This behavior is not unexpected for stochastic diffusion training because each loss measurement reflects the sampled training batch, diffusion timestep, and injected noise rather than a deterministic full-dataset objective.

Training Progress

A simplified view of the run is shown below.

Training stage Approximate loss behavior
Epochs 0–4 Rapid initial reduction from 0.358 to ~0.017
Epochs 5–15 Mostly low losses with occasional spikes
Epochs 16–30 Stable low-loss regime with stochastic fluctuations
Epochs 31–40 Continued low-loss optimization as LR decayed
Epochs 41–49 Very low learning rate and final convergence phase

The final epoch completed at:

Epoch: 49
Step: 3,149
Reported loss: 0.0384
Learning rate: 0

Because the logged value represents the final observed batch rather than an epoch-averaged validation metric, it should not be interpreted as the model's overall evaluation score.

Training Throughput

Training throughput was generally stable at approximately:

7.1 to 7.3 iterations per second

Most epochs contained:

63 / 63 training iterations

Periodic generation loops ran at approximately:

34.8 iterations per second

This suggests relatively stable computational throughput during the training run.

Results

Training completed successfully for all 50 epochs and 3,149 optimization steps.

The logs show:

  • rapid reduction in loss during early training
  • stable single-GPU throughput
  • successful completion of the full training schedule
  • gradual learning-rate warm-up followed by decay
  • periodic image-generation passes throughout training
  • occasional stochastic loss spikes despite a generally low-loss regime

These results demonstrate successful end-to-end execution of the diffusion training pipeline.

They do not, by themselves, establish image-generation quality. Diffusion training loss is useful for monitoring optimization but should be interpreted alongside generated samples and dedicated generative evaluation metrics.

Model Examination

One important observation from the training run is that diffusion loss is visibly noisy.

For example, the reported loss moved from:

0.0151 at epoch 6

to:

0.129 at epoch 7

and subsequently returned to:

0.0134 at epoch 8.

Similar short-term increases occurred later in training.

This illustrates an important property of diffusion-model optimization: individual training loss measurements can fluctuate substantially because the training objective depends on randomly sampled images, noise levels, timesteps, and noise realizations.

As a result, isolated loss values should not be used as the sole criterion for judging generative quality.

How to Get Started

The exact loading code depends on how the model components were serialized in this repository.

For a standard Hugging Face Diffusers pipeline, inference typically follows this pattern:

import torch
from diffusers import DiffusionPipeline

model_id = "pynk17/my-small_diff-model"

pipe = DiffusionPipeline.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
)

pipe = pipe.to("cuda")

prompt = "your prompt here"

image = pipe(prompt).images[0]
image.save("generated_image.png")
Downloads last month
43
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train pynk17/my-small_diff-model