Sxela / IP-Adapter

The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.

Geek Repo:Geek Repo

Github PK Tool:Github PK Tool

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models


Introduction

we present IP-Adapter, an effective and lightweight adapter to achieve image prompt capability for the pretrained text-to-image diffusion models. An IP-Adapter with only 22M parameters can achieve comparable or even better performance to a fine-tuned image prompt model. IP-Adapter can be generalized not only to other custom models fine-tuned from the same base model, but also to controllable generation using existing controllable tools. Moreover, the image prompt can also work well with the text prompt to accomplish multimodal image generation.

arch

Release

  • [2023/8/18] 🔥 Add code and models for SDXL 1.0. Demo is here.
  • [2023/8/16] 🔥 We release the code and models.

Dependencies

  • diffusers >= 0.19.3

Download Models

you can download models from here. To run the demo, you should also download the following models:

How to Use

  • ip_adapter_demo: image variations, image-to-image, and inpainting with image prompt.

image variations

image-to-image

inpainting

structural_cond

multi_prompts

Best Practice

  • If you only use the image prompt, you can set the scale=1.0 and text_prompt=""(or some generic text prompts, e.g. "best quality", you can also use any negative text prompt). If you lower the scale, more diverse images can be generated, but they may not be as consistent with the image prompt.
  • For multimodal prompts, you can adjust the scale to get best results. In most cases, setting scale=0.5 can get good results. For the version of SD 1.5, we recommend using community models to generate good images.

Citation

If you find IP-Adapter useful for your your research and applications, please cite using this BibTeX:

@article{ye2023ip-adapter,
  title={IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models},
  author={Ye, Hu and Zhang, Jun and Liu, Sibo and Han, Xiao and Yang, Wei},
  booktitle={arXiv preprint arxiv:2308.06721},
  year={2023}
}

About

The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.

License:Apache License 2.0


Languages

Language:Jupyter Notebook 99.8%Language:Python 0.2%