# How to Use an AI Image Generator

Learn SeaArt AI with our expert AI image generator guides.

## **Welcome to SeaArt AI Image Generator Guide!**

SeaArt is a premier free AI art generator. Dive into a thriving AI content community and explore over 1000k+ models and LoRas. From art to illustrations and paintings, SeaArt effortlessly enhances your creativity.

Start for free today and elevate your creative workflow!

**SeaArt APP:** [*https://seaart-all.onelink.me/qoIB*](https://seaart-all.onelink.me/qoIB/DC)

**Discord:** <https://discord.gg/HA5TbrsN7s>

> #### You can contact us via our Discord server if you have any questions.

## <mark style="background-color:yellow;">Follow us on Socials:</mark>

* Instagram: [@seaartai](https://www.instagram.com/seaartai/)
* Twitter: [@SeaArt\_Ai](https://x.com/SeaArt_Ai)
* TikTok: [@seaart.ai](https://www.tiktok.com/@seaart.ai?lang=en)
* Youtube: [@SeaArt](https://www.youtube.com/channel/UC-hWsHj8797Sv2sf20-6-3g)
* Facebook:[ @SeaArt\_Ai](https://www.facebook.com/SeaArtAiOfficial)
* Reddit: [@SeaArt\_Ai](https://www.reddit.com/r/SeaArt_Ai/)

## **🐾Starting** [**SeaArt.AI**](http://seaart.ai/) **!!!**

## What is SeaArt.AI?

[SeaArt.AI](https://www.seaart.ai/) is a globally accessible, free online AI art creation platform. Wherever you are, as long as you have an internet connection, you can access SeaArt.AI's services by visiting our website at [seaart.ai](https://www.seaart.ai/). Whether you're on a computer, Android device, or Apple device, our platform offers a convenient and engaging art creation experience.<br>


# 1-SeaArt AI Basic Page

Learn the basics of SeaArt AI, and start your journey of AI art creation today.

## New User Registration

**STEP1:** Click on <mark style="background-color:yellow;">'R</mark>egister<mark style="background-color:yellow;">'</mark> at the right side of the top menu bar to open the login interface.

<figure><img src="/files/SJholkHH0pV6GhflLvWH" alt=""><figcaption></figcaption></figure>

**STEP2:** Choose the appropriate method to log in or register.

## &#x20;**Registration and Login Guide**

**1. Email Verification Steps and Common Issues (What to do if you don't receive the verification email?)**

● Check whether the email address entered is correct.

● Check the spam/junk/promotions folder.

● Wait a few minutes and refresh your inbox.

● If still not received, try using a popular email provider (e.g., Gmail, Outlook).

● If multiple attempts fail, please contact customer service for assistance.

**2. How to log in using a third-party account (e.g., Google)**

On the login page, select "Log in with Google" and follow the prompts to authorize. If login fails, try clearing your browser cache, switching browsers, or changing network environments. If the issue persists, contact customer service with the specific error message.

**3. Basic Troubleshooting for Third-Party Login Failures**

● Clear browser cache and cookies.

● Try another browser or incognito mode.

● Check if your Google account is functioning normally.

● If it still doesn't work, take a screenshot of the error and contact customer service.

**4. Forgot Password / Change Password**

Click "Forgot Password" on the login page, enter your registered email, and the system will send a password reset link. Follow the instructions in the email to set a new password. If you don't receive the email, refer to the "Email Verification" section above.

<figure><img src="/files/pNeVhsBf6bCVnOQpRHO9" alt=""><figcaption></figcaption></figure>

## **Quick Start for Beginners**

**Steps to Create Your First Work**

Click Create → Select Image/Video → Choose Model → Enter Prompt → Click Generate

<figure><img src="/files/FQ8vOfxeV4cW03zJsBGK" alt=""><figcaption></figcaption></figure>

## Page Introduction

### Homepage Layout

The navigation bar on the left allows you to explore various features.

Scroll through the window to view limited-time events and new feature releases.

The tabs at the top make it easier for you to view categorized images.AI art

> **Page Entry:** Click "Generate" in the top right corner to access image generation, AI character, audio, canvas, and workflow features.

#### **Classic Mode**

Classic Mode supports [<mark style="background-color:yellow;">Text to Image</mark>](/guide-1/2-seaart-ai-basic-function/2-1-text-to-image)<mark style="background-color:yellow;">,</mark> [<mark style="background-color:yellow;">Image to Image</mark>](/guide-1/2-seaart-ai-basic-function/2-2-image-to-image)<mark style="background-color:yellow;">, and</mark> [<mark style="background-color:yellow;">ControINet</mark>](/guide-1/2-seaart-ai-basic-function/2-3-controlnet)<mark style="background-color:yellow;">,</mark> while also offering some AI creative <mark style="background-color:yellow;">tools.</mark>

#### AI Tools

* **Animate:** Convert static images into videos

<figure><img src="/files/mmiIMUsF6wkEezy6ZG0E" alt="A bird on a branch" width="240"><figcaption></figcaption></figure>

* **Expansion:** After uploading the image, drag to expand the frame, select a Model with the same style as the original image, and only input the prompts related to the part of the image you wish to expand.
* **Face Swap:** After uploading the original image, add the face that needs to be replaced on the right side.
* **Upscale:** Increase the clarity of the image without changing the details.
* **Character Repair:** Used to repair faces, hands, and bodies in images, it is recommended to choose a Model with the same style as the original image.

<figure><img src="/files/U0y56ZiTniS3QrEGYgvR" alt="Before and after comparison of image repairing"><figcaption></figcaption></figure>

* **Describe:** After uploading the image, display the prompts.
* **Preview:** After uploading the finished image, select ControlNet to display the 'original image'.

### Explore

**Page Entry:** Click <mark style="background-color:yellow;">'Home'</mark> in the upper left corner.

**Tags:** Display images of a specific tag category.

**Search Bar:** Choose to search for <mark style="background-color:yellow;">Works/Model/Tags/User/Post/Workflow/AI Apps/Canvas/AI Characters.</mark>

**Online Customer Service:** Click the customer service button in the lower right corner to contact online customer service. There may be a queue, so please be patient and wait a moment\~

### Work Details

* **Page Entry:** Click on the corresponding work or hover the mouse over it.
  * On the work details page, you can see the prompts used for creating the image, the model, other detailed information, and user comments.
  * Supports high-definition restoration of the image, variation (V), and background removal based on this.
  * Variations (V): Enter the Img2Img page to modify image details.

Click to report the work.

<figure><img src="/files/KXkdZdwn8uvlNPBMJ0NU" alt=""><figcaption></figcaption></figure>

Click to obtain the link/image for sharing the work.

<figure><img src="/files/DArOQXCfO2d5wsjX4NVZ" alt=""><figcaption></figcaption></figure>

Click to download the work.

<figure><img src="/files/eQgMq0wVYKIjLMrOF6c6" alt=""><figcaption></figcaption></figure>

### Personal

<mark style="background-color:red;">Hover the mouse over the avatar:</mark>

**Green Mode:** Activating Green Mode can hide images that are not suitable for display, including those with gore, violence, pornography, etc.

**Language:** Change the website language

<figure><img src="/files/xZxOV6dnpWW4upzJEzav" alt=""><figcaption></figcaption></figure>

**Settings:**

**Personal Settings:**

Change name

Bookmark website tags

<figure><img src="/files/DAJfdjJ3rFMEAo6dxGq1" alt=""><figcaption></figcaption></figure>

**System Settings:**

Whether to publicly display your own creations

Whether to turn off auto-translation of prompt words - refers to whether the prompt words entered in the creation flow are automatically translated

<figure><img src="/files/CaKAWwrYyaCKiEB4GzmR" alt=""><figcaption></figcaption></figure>

<mark style="background-color:red;">Click on the avatar in the upper right corner to enter the personal center:</mark>

View the **User ID** and personal invitation code in the upper left corner.

**Works:** View the images you have created.

Click "**Manage**" in the lower right corner:

Choose to post works.

Choose to save works to a created folder.

Find works by time.

**Favorites:** View favorited works/Model/Post/canvas/WorkflowI.

**Post:** View posted works.

**Canvas:** View the canvases you have created.

**Creator Profit Center:** You can receive cash rewards;&#x20;

click here for specific rules.

{% content-ref url="/pages/MvgHDJQdi9shrivEb3iU" %}
[SeaArt.AI Creator Incentive Program](/guide-1/6-permanent-events/seaart.ai-creator-incentive-program)
{% endcontent-ref %}

### Tasks

Click on "Tasks" on the left to enter the Task Center, where you can find daily, weekly, and one-time tasks.

<figure><img src="/files/2Hywwl97RtYW3Xp74nYf" alt=""><figcaption></figcaption></figure>

In the entire record, you can view the Total Earnings and Total Spending of stamina/Credits.

<figure><img src="/files/YiMwkpVNgteWuwRoIWo0" alt=""><figcaption></figcaption></figure>

### SeaArt Mall

**Page Entry:** Click on the VIP icon at the top right to enter the SeaArt Mall.

**Introduction to SeaArt Mall:**

**Stamina:** Refreshes daily. Basic users receive 150 Stamina Per Day, while VIP users can obtain more.

**Credits:** Never expires. During creation, stamina will be used first, If stamina is insufficient, Credits will be used.

**VIP Privileges:** VIPs can create unlimitedly for free. For more benefits, See the image below.

**Credits Recharge:** Credits can be used for Model training, AI creation, and paid model usage, and it never expires.

<figure><img src="/files/TZVjFk8QV3vbVKFjMDrZ" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/NsWRAp8zgtYcOhYHlJ0f" alt=""><figcaption><p>The exact price is based on the website.</p></figcaption></figure>

#### Credits

Credit will not expire or be reset and can be purchased through the SeaArt Mall. Meanwhile, Credit can also be obtained by completing tasks and participating in Community or Website events.

**a. Credit can be purchased through the SeaArt Mall.**

Tap the "VIP" icon in the top right corner of the page to check the price and buy Credits in the SeaArt Mall.

**b. Get Credits by completing tasks**&#x20;

Tap Tasks in the bottom right corner to check the Task List. Complete the corresponding tasks to earn the rewards.

**c. Participate in Community or Website events.**&#x20;

We will periodically hold events on our website and in the Discord. Please pay attention to the Community Notice [#announcement](https://discord.com/channels/1089843669944778822/1089854034866884609). Don't miss the opportunity to get rewards!

#### Stamina

The free users can get 150 Stamina daily, VIP users can receive 300/700/2100/3500 stamina daily for free, depending on their level.

Your Stamina will be reset at 0:00 UTC daily. Hover the mouse on your avatar in the top right corner to see how much Stamina is left.

Stamina can be used for your daily creations. You can check it here.

<figure><img src="/files/XwoQipdpszxKuyLfm7KH" alt=""><figcaption></figcaption></figure>

### Contact Us

**Page Entry:** Click on the relevant icons to join our official community and social media channels.

Our staff will be online to answer your questions, and you'll be the first to receive the latest updates about SeaArt. Engage with AI experts and exchange insights on AI technology!


# 2-SeaArt AI Basic Function

Explore these basic functions of SeaArt, and learn how to use them to create stunning AI visual content.


# 2-1 Text to Image

Learn what is text to image and how to use AI text-to-image generator with easy, step-by-step instructions.

> **Have you ever encountered similar problems while using SeaArt: the control effect is not ideal, or the drawing results do not reflect the added prompts, and so on? In this article, we will comprehensively introduce how to master the operation of text-to-image and grasp the strategy of writing efficient prompt words.**

### What is Text to Image?

In SeaArt AI, there are three modes of drawing: <mark style="background-color:yellow;">Text to Image,</mark> [<mark style="background-color:yellow;">Image to Image</mark>](/guide-1/2-seaart-ai-basic-function/2-2-image-to-image)**,** and [<mark style="background-color:yellow;">ControINet</mark>](/guide-1/2-seaart-ai-basic-function/2-3-controlnet)

The basic steps for drawing are: <mark style="background-color:red;">select a model→ enter prompts→ set parameters→ generate.</mark>

The model determines the <mark style="background-color:yellow;">style</mark>, prompts define the <mark style="background-color:yellow;">content of the image</mark>, and parameters refine the <mark style="background-color:yellow;">preset characteristics of the image.</mark>

## **Prompt Writing – Basics and Advanced**

**What is a Prompt? How to Write a Good Prompt?**

A prompt is a textual description that guides the AI to generate content. A good prompt should be clear and specific, covering key aspects like style, content, and details. Example: "Japanese anime style, girl, under a cherry blossom tree, smiling, sunny day." We recommend referencing community examples and gradually building experience.

**Basic Prompt Syntax and Structure**

● Use commas to separate keywords; word order affects results.

● Use parentheses or weights to emphasize keywords, e.g., (cat:1.2).

● Multi-language support is available, but English or the platform-recommended language is preferred.

● Use negative prompts to exclude unwanted elements.

**How to Set Prompt Weights (Detailed Explanation)**

Weighting emphasizes keyword importance. Common format: (keyword:1.5) means weight = 1.5 (greater influence).

Multiple keywords can have individual weights, e.g., (cat:1.2), (dog:0.8).

Support varies slightly across models; please refer to the docs or community tips.

**How to Use Negative Prompts**

In the negative prompt field, input undesired elements like: "blurry, low quality, watermark." Use weights to strengthen exclusions, e.g., (multiple people:1.5), to control for solo portraits.

Negative Prompts are particularly useful when some Models have a poor understanding of specific details (e.g., hand structures), as they help avoid these elements and improve image quality.

For example, include: (bad hands, bad anatomy, bad body, bad face, bad teeth, bad arms, bad legs, deformities: 1.3)

Input Prompts: natural language/phrase form&#x20;

**Natural language:** A girl with black hair dancing&#x20;

**Phrase form:** A girl, black hair, dancing

The role of prompts is to guide and assist the model in the drawing process, rather than being a strict requirement. Even if your input is just a casual sentence, the model can still create an image for you, and the result might even be quite good.

<mark style="color:red;">\*Rich prompts can better control the final output effect. In the later fine-tuning process, specific keywords can be quickly modified and verified for their impact on the drawing result.</mark>

### Universal Prompt Formula

An effective prompt is like assigning a task to the AI art generator. If the instruction is vague, such as merely saying "design a picture" without specifying elements and purpose, the result is often unpredictable. Therefore, detailed and specific instructions can greatly improve the quality and relevance of the outcome.

For example, if the prompt simply inputs "a girl," it does not mention the girl's attire, scene, camera angle, etc., and the AI can only perform based on the model's historical experience during training. Thanks to the model's capabilities, the drawing results we get are still quite good. However, if there are specific requirements for the content of the screen, such efficiency is very low.

When we add other descriptive words for content, the image will become much more stable.

An ideal prompt formula includes elements such as the <mark style="background-color:yellow;">main content, environmental background, composition, image settings, and reference style</mark>, each affecting the drawing outcome to different extents.

<mark style="color:red;">\*This formula is a reference, not a strict rule for every prompt creation. First, determine the main content's impact, then optimize details according to personal needs.</mark>

<figure><img src="/files/Na61ixdBMm35xJssfICV" alt=""><figcaption></figcaption></figure>

1. **Main content:** Main content describes the primary subject, like <mark style="background-color:yellow;">people or animals, their clothing, expressions, fur, actions, or the material of objects</mark><mark style="color:blue;">.</mark> Generating multiple subjects together might pose issues; it's advisable to create each subject separately and then use ControINet generation for integration.
2. **Environmental background:** Environmental background sets the scene and auxiliary elements like <mark style="background-color:yellow;">sky color, surroundings, lighting, and color tone,</mark> enhancing the image's atmosphere and highlighting its theme.
3. **Composition of shots:** Composition adjusts the camera angle and perspective, such as <mark style="background-color:yellow;">depth of field emphasis or object layout,</mark> significantly boosting the visual impact.
4. **Image settings:** Image settings include terms to enhance visual expressiveness, like <mark style="background-color:yellow;">detail richness, photography quality, and cinematic effect</mark><mark style="color:blue;">.</mark> Image resolution and detail level are mainly determined by size, with post-processing techniques like Upscale further enhancing details.
5. **Reference style:** Describes the desired artistic style and mood, such as mentioning an <mark style="background-color:yellow;">artist's name, art techniques, era, or colors.</mark> However, the image style is largely determined by the model; if the model hasn't been trained on specific artistic style keywords, it might not understand them. For specific style requirements, using a model trained in that style might yield better results than merely using prompts.

**Remix:** If you find writing prompts too complex, you can look for inspiration from the AI-generated images on the homepage and use the one-click reuse of existing parameters and prompt words to simplify the creation process.

### Emphasized Prompts

Emphasizing prompts relies on parentheses and numerical values to control the weight of specific prompts. The higher the weight value, the more the model prioritizes that prompt, focusing on rendering that part during the process. As a result, the final image will reflect more of the corresponding information. Conversely, less emphasis will result in less representation of that content in the image.

One method is to <mark style="background-color:yellow;">increase weight through the use of parentheses</mark>, and the other is to directly <mark style="background-color:yellow;">enter numerical values</mark>, with the latter being the more commonly used approach.

There are three types of parentheses for controlling the weight of prompt words:

* Round parentheses ( ): Each layer increases the original weight by 1.1 times.
* Square brackets \[ ]: Each layer decreases the original weight to 0.9 times.
* Moreover, parentheses support multiple layers of stacking, with each layer representing a weight multiplied by a fixed factor.

<figure><img src="/files/dxsMoOFz36gm2Hkdttmy" alt="Emphasize prompts by adding parentheses"><figcaption></figcaption></figure>

For example, by default, the girl's clothes will be a combination of yellow and orange. However, when "**(((orange coat)))**" is used, with the parentheses indicating an increase in emphasis, the model's depiction of the orange coat is enhanced, resulting in <mark style="background-color:yellow;">more orange</mark> appearing in the coat in the final image.

<figure><img src="/files/NWyhzIbCqYp2PLIrlQPK" alt="Before and after comparison of emphasized prompts" width="563"><figcaption></figcaption></figure>

Conversely, when "**\[\[orange coat]]**" is used, with the square brackets indicating a decrease in emphasis, the orange elements are diminished. The model will then prioritize the remaining keywords "**((Yellow coat))**", leading to the coat appearing <mark style="background-color:yellow;">more yellow</mark> in the final image.

<figure><img src="/files/A3rNQ6Vsv9dz9wy8ohXn" alt="Before and after comparison of AI images" width="563"><figcaption></figcaption></figure>

Directly input numerical values to control weight.

for example, by default, the hair is presented in green and red colors. If we set the weight after "(green hair)" to 0.9, it means the weight of the green hair part is reduced to 0.9 times its original value. Similarly, if we want to increase the weight of the green hair, we can simply enter 1.1 afterward.

<figure><img src="/files/CZbHBngPvGSOjEx1CsGL" alt="Three examples of AI anime images with different prompts"><figcaption></figcaption></figure>

<mark style="color:red;">\*Although the emphasis on keywords' weight can vary from 0.1 to 100, considering the potential effect deviations caused by extreme weight values, it is recommended to keep the weight between 0.5 and 1.5 for optimal image results.</mark>

<mark style="background-color:red;">For specific parameter settings, click here to view details.</mark>

{% content-ref url="/pages/rNAxYQXt8yBJlPmXAKfD" %}
[4-Parameters](/guide-1/4-parameters)
{% endcontent-ref %}


# 2-2 Image to Image

Ready to transform images? Step into the world of image-to-image and learn its parameters and workflow.

> **In the practical application of AI painting, due to the uncertainty of the initial images produced by the model, the actual controllability of the output images is not high. Subsequently, we can use the "Image to Image" function to modify the pictures towards our desired direction, thereby enhancing the controllability of the generated images.**

## What is Image to Image?

The "Image to Image" function is an AI-based image generation technology that allows users to generate new images based on an existing picture combined with text descriptions. This technology is significant because it can create new visual content by mixing image and text prompts according to the specific needs of users.

Simply put, the process of considering <mark style="background-color:yellow;">both the prompt words and the image information</mark> in the reference picture and then drawing is what constitutes the image-to-image generation.

## Analysis of Img to Img Parameters

**Intelligent Analysis**

<mark style="background-color:yellow;">Automatically inferring matching</mark> prompts based on the provided image, as well as a model that matches the image. However, the prompts generated by intelligent analysis may contain incorrect prompts, so it is recommended to manually perform a <mark style="background-color:yellow;">second screening</mark>. This function primarily serves as a reference for writing prompt words.

<figure><img src="/files/BO3a3PDFV8oidE9AXm4q" alt="SeaArt intelligent analysis feature"><figcaption></figcaption></figure>

**Workflow of Img to Img**

<mark style="background-color:red;">Workflow: Upload Reference Image - Set Model Prompts - Set Parameters - Generate</mark>

After uploading the reference image, open intelligent analysis, which will automatically fill in the prompts, Model, and image size. It is recommended to readjust the prompts according to actual needs. The parameter settings are the same as for generating initial images. Finally, click to generate the image, and AI will create a new image based on the reference picture and the user's instructions.

<mark style="color:red;">\*The larger the redrawing extent, the greater the difference from the original image. It is generally set between 0.4-0.8.</mark>

**Denoising Strength:** This parameter controls the degree of divergence in the redrawing process based on the original image. The higher the value, the more freedom the model has during the redrawing process, and the <mark style="background-color:yellow;">greater the difference</mark> between the drawing result and the original reference image.

When the Denoising Strength is too high, it becomes difficult to associate the drawn image content with the original image, hence, we typically keep the value of the redrawing extent <mark style="color:red;">between 0.4 and 0.8.</mark>

**Partial Repainting**

Partial Repainting allows for modifications and adjustments to <mark style="background-color:yellow;">specific areas within an image.</mark> This feature is particularly suited for fine-tuning local details. It also requires the addition of prompts to guide the modification content. This is used when the majority of the image content is satisfactory, but there is a need to adjust some detail elements.

After uploading the image, click the brush on the right to enter the Partial Repainting area. Then, you can smear the image. After smearing, fill in the prompts for the smeared area in the prompt box.

After using Partial Repainting, only the selected area has been redrawn, while the other areas remain unchanged.


# 2-3 ControlNet

Master AI image generation with ControlNet. Learn about its preprocessors, how it works, and how to use it to create stunning AI art.

## What is ControINet?

ControlNet is a plugin used for controlling AI image generation. It employs a technology known as "Conditional Generative Adversarial Networks" (CGANs) to generate images. Unlike traditional Generative Adversarial Networks, ControlNet allows users to finely control the generated images, such as uploading line drawings for AI to colorize, or controlling the posture of characters, generating image line drawings, etc.

Different from traditional drawing models, a complete ControlNet consists of two parts: a <mark style="background-color:yellow;">Preprocessing Model</mark> and a <mark style="background-color:yellow;">ControlNet.</mark>

* Preprocessing Model: Responsible for extracting the spatial semantic information from the original image and converting it into a visual preview image, such as line drawings, depth maps, etc.
* ControlNet: Processes more fundamental structured information like lines and depth of field.

### **Canny**

**Basic Information**

The Canny model primarily identifies <mark style="background-color:yellow;">edge information</mark> in input images, capable of extracting precise line drawings from uploaded pictures. It then generates new scenes consistent with the original image's composition based on specified prompts.

<figure><img src="/files/KFSxtlRsFhHP1x4JxdaS" alt="Before and after comparison of using the Canny model to process a 3d cartoon image" width="563"><figcaption><p>Original / Preprocessing</p></figcaption></figure>

Preprocessor:

* **canny:** Hard edge detection.
* **invert:** Inverts the colors to black lines on a white background, reversing the colors of the line drawings.

<figure><img src="/files/kEWN0hDjhHSvGER4onU6" alt="Before and after comparison of using the Canny model to process a sketch image" width="563"><figcaption><p>Original / invert</p></figcaption></figure>

invert is not unique to Canny and can be used in conjunction with most line drawing models. When we select ControINet types like Line Art or MLSD t recognition, the invert is available.

**Method of Operation**

Operational Sequence:

<mark style="background-color:yellow;">Upload Image - Select Model - Select ControINet Type - Enter Prompts - Generate</mark>

Intelligent Analysis:

Reverse-infer the image prompts and model. If a different style from the original image is desired, it is recommended to turn off intelligent analysis.

Parameter Settings:

<mark style="background-color:yellow;">Preprocessing Resolution</mark>

The preprocessing resolution affects the output resolution of the preview image. Since the aspect ratio of the image is fixed, and the default output is a 1x image, the resolution setting essentially determines the horizontal size of the preview image. For example, if the original image size and the target image size are both 512x768, when we set the preprocessing resolution to 128, 256, 512, 1024, the preprocessed image size will change to 128x192 (0.25x original), 256x384 (0.5x original), 512x768 (original), and 1024x1536 (2x original), respectively.

<mark style="color:red;">In general, the higher the resolution setting, the richer the details of the generated image.</mark>&#x20;

<mark style="color:red;">\*</mark>Sometimes, when the preprocessing detection image and the final <mark style="background-color:yellow;">image size are inconsistent</mark>, it can lead to damaged drawn images, with clear pixelation at the edges of figures in the final drawing.

<figure><img src="/files/5ULODY6n2xYlYaShduT0" alt="Examples of images processed by the Canny model with different resolution"><figcaption></figcaption></figure>

<mark style="background-color:yellow;">Control Weight</mark>

Determines the strength of the ControINet. The higher the intensity, the more pronounced the control over the image effect, and the closer the generated image is to the original.

<figure><img src="/files/TcJu8Mi5RmXahoyAdxZf" alt="Examples of superman images processed by the Canny model with different weight" width="506"><figcaption></figcaption></figure>

<mark style="background-color:yellow;">Control Mode</mark>

Used to switch the weight proportion between ControlNet and the prompt words. The default setting is balanced.

Prioritize Prompts: The control diagram effect will be weakened.

Prioritize Pre-processing Image: The control diagram effect will be enhanced.

<mark style="background-color:yellow;">Generation Results</mark>

From the generated results, it can be seen that the basic composition is exactly the same as the original image, but the details are completely different. If you need images with other changes, such as hair color, facial details, clothing, etc., you can adjust the keywords and parameters to achieve the desired effects.

<figure><img src="/files/Dxz3BDz6o4f8ArZDmzPF" alt="Four examples of AI-generated superman images with different models"><figcaption></figcaption></figure>

### **OpenPose Full**

**Basic Information**

OpenPose Full can achieve precise control over <mark style="background-color:yellow;">human body movements and facial expression features.</mark> It's capable not only of generating poses for a single person but also for multiple people.

OpenPose Full can <mark style="background-color:yellow;">identify key structural points of the human body</mark> such as the head, shoulders, elbows, knees, etc., while ignoring details of clothing, hairstyles, and backgrounds, ensuring the true reproduction of poses and expressions.

**Preprocessor**

Human Pose Recognition

The default processors are from the openpose series, including <mark style="background-color:yellow;">openpose, face, faceonly, full, hand.</mark> These five preprocessors are used to detect facial features, limbs, hands, and other human body structures respectively.

Animal Pose Recognition

It is recommended to use the animal\_openpose processor, which can be used in conjunction with specialized preprocessing models, such as control\_sd15\_animal\_openpose\_fp16.

<figure><img src="/files/Z2AiFjAod08HvmvwSydc" alt="Animal pose recognition" width="563"><figcaption></figcaption></figure>

In general, using the default openpose\_full preprocessor is sufficient.

### **Line Art**

**Basic Information**

Line Art is also about extracting the edge line art from images, but its use cases are more specific, including two directions: realistic and anime.

**Preprocessor**

Line Art

More suitable for realistic images, the extracted line art is more restorative, retaining more edge details during detection, thus the control effect is more significant.

Difference between Line Art and Canny

Canny: Hard straight lines, uniform thickness.

Line Art: Obvious brushstroke traces, similar to real hand-drawn drafts, allowing clear observation of thickness transition under different edges.

Line Art retains <mark style="background-color:yellow;">more details,</mark> resulting in a <mark style="background-color:yellow;">relatively softer</mark> image, and is more suitable for line art coloring functions.

Canny are <mark style="background-color:yellow;">more precise and simplify</mark> the content of the image.

<mark style="color:red;">\*Line Art can be used for coloring draft images, fully following the draft.</mark>

### **Depth**

**Basic Information**

Depth, also known as distance images, intuitively reflects the three-dimensional depth information of objects in a scene. Depth is displayed in black and white; <mark style="background-color:yellow;">the closer an object is to the camera, the lighter (whiter) its color; conversely, the farther away it is, the darker (blacker) its color.</mark>

<figure><img src="/files/c7oRRtHAMvAA3AUjt9hZ" alt="Comparison of the original image and the image processed with Depth"><figcaption><p>Original / Depth</p></figcaption></figure>

Depth can extract the <mark style="background-color:yellow;">foreground and background</mark> relationship of objects from an image, create a depth map, and apply it to image drawing. Therefore, when it is necessary to <mark style="background-color:yellow;">clarify the hierarchical relationship of objects in a scene,</mark> depth detection can serve as a powerful auxiliary tool.

It is recommended to use the <mark style="background-color:red;">depth\_midas</mark> preprocessor to achieve better image output results.

<figure><img src="/files/o058oguiDENmrnIa1s7M" alt="Comparison of the original image, depth image and result image"><figcaption><p>Original / Depth / Result</p></figcaption></figure>

### **Normal Bae**

**Basic Information**

Normal Bae involves generating a normal map based on the light and shadow information in the scene, <mark style="background-color:yellow;">thereby simulating the details of the object's surface</mark> texture and accurately restoring the layout of the scene's content. Therefore, model recognition is often used to reflect more <mark style="background-color:yellow;">realistic light</mark> and shadow details on object surfaces. In the example below, you can see a significant improvement in the lighting and shadow effects of the scene after drawing with model recognition.

When using, it is recommended to select the <mark style="background-color:red;">normal\_bae</mark> preprocessor for a more noticeable improvement in lighting and shadow effects.

<figure><img src="/files/gEBzcfygL6qK5gQ3rsfW" alt="Comparison of the original image, Normal Bae image and result image"><figcaption><p>Original / Normal Bae / Result</p></figcaption></figure>

### **Segmentation**

**Basic Information**

Segmentation can divide the scene into different blocks while detecting content outlines, and assign semantic annotations to these blocks, thereby achieving <mark style="background-color:yellow;">more precise control over the image.</mark>

Observing the image below, we can see that the image after semantic segmentation detection includes different colored blocks. Different contents in the scene are assigned different colors, such as characters marked in red, the ground in brown, signboards in pink, etc. When generating images, the model will produce specific objects within the corresponding color block range, thus achieving more accurate content restoration.

When using, it is recommended to select the default <mark style="background-color:red;">seg\_ufade20k</mark> preprocessor. Users can also modify the image content by filling in color blocks in the preprocessing image.

<figure><img src="/files/wshAJxryKzWXC1XUAhN2" alt="Comparison of the original image, Segmentation image and result image"><figcaption><p>Original / Segmentation / Result</p></figcaption></figure>

### **Tile Resample**

Tile Resample can convert low-resolution images into higher-resolution versions while minimizing quality loss.

Three preprocessors: tile\_resample, tile\_colorfix, and tile\_colorfixsharp.

\*In comparison, the default resample offers more flexibility in drawing, and the content will not differ significantly from the original image.

<figure><img src="/files/4CHyTAbOOZrBSuvyqcgB" alt="Examples of AI cartoon wolf images processed with different preprocessors"><figcaption></figcaption></figure>

<mark style="color:red;">\*</mark>In comparison, the default <mark style="background-color:red;">resample</mark> offers more flexibility in drawing, and the content will not differ significantly from the original image.

### **MLSD**

**Basic Information**

MLSD recognition extracts <mark style="background-color:yellow;">straight edge</mark> lines from the scene, making it particularly useful for delineating the linear geometric boundaries of objects. The most typical applications are in the <mark style="background-color:yellow;">fields of geometric architecture, interior design, and similar areas.</mark>

<figure><img src="/files/bTSFdrIQHa0HysbM7qNl" alt="Comparison of the original image, MLSD recognition image, and result image"><figcaption></figcaption></figure>

### **Scribble HED**

**Basic Information**

Scribble HED resembles crayon scribble line drawings, offering more freedom in controlling the image effect.

Preprocessors: HED, PiDiNet, XDoG, and t2ia\_sketch\_pidi.

As can be seen from the images below, the first two preprocessors produce thicker outlines that are more in line with the hand-drawn effect of doodles, while the latter two produce finer lines, suitable for realistic styles.

<mark style="color:red;">\*Can be used for coloring draft images, with a certain degree of randomness.</mark>

**HED**

**Basic Information**

HED creates <mark style="background-color:yellow;">clear and precise boundaries around objects,</mark> with an output similar to Canny. Its effectiveness lies in the ability to <mark style="background-color:yellow;">capture complex details and contours</mark> while retaining detailed features (facial expressions, hair, fingers, etc.). The HED preprocessor can be used to modify the style and color of an image.

Compared to Canny, HED produces <mark style="background-color:yellow;">softer lines and retains more details.</mark> Users can choose the appropriate preprocessor based on their actual needs.

### **color\_grid**

**Basic Information**

Through the use of preprocessors, we can obtain results from color block processing, where the generated images will be redrawn based on the original colors.

### **shuffle**

**Basic Information**

By randomly shuffling all information features of the reference image and then recombining them, the generated image may differ from the original in structure, content, etc., but a hint of stylistic correlation can still be observed.

The use of content recombination is not widespread due to its relatively poor control stability. However, using it to gain inspiration could be a good choice.

<figure><img src="/files/RZHRjtImix05VHEaxaF0" alt="Comparison of original vs shuffle vs result" width="563"><figcaption></figcaption></figure>

### **Reference Generation**

**Basic Information**

To generate a new image based on the reference original, it is recommended to use the default <mark style="background-color:red;">"only"</mark> preprocessor.

**Control Weight:** The higher the value, the stronger the stability of the image, and the more obvious the traces of the original image's style will be preserved.

<figure><img src="/files/tvU7DoJmYoFn6knKXBA3" alt="Different reference images generated based on the cartoon fox image"><figcaption></figcaption></figure>

### **recolor**

Filling in colors for images is very suitable for repairing some black-and-white old photos. However, it cannot guarantee that colors appear accurately in specific positions, and there might be cases of color contamination.

**Preprocessors:** "intensity" and "luminance", with <mark style="background-color:red;">"luminance"</mark> is recommended.

<figure><img src="/files/n1SCfrZYueksJ5kDFHYO" alt="Comparison - original vs recolor_luminance vs recolor_intensity"><figcaption></figcaption></figure>

### **ip\_adapter**

**Basic Information**

Turning an uploaded image into image prompts allows it to recognize the artistic style and content of the reference image, and then generate similar works. It can also be used in conjunction with other ControINet.

<figure><img src="/files/z6LaQ0OyyHtRldeM7S3L" alt="apply ip_adapter to an image to genrated a new one"><figcaption></figcaption></figure>

**Method of Operation**

1. Upload the original image A that needs to be generated, and select ControINet options like Canny, openpose, Depth, etc.
2. Add a new ControINet, ip\_adapter, upload the style image B you want to inherit, and finally click Generate.

<mark style="background-color:red;">Result: Image A with the style of B.</mark>


# 2-4 AI Apps

Explore practical and fun AI Apps for transforming styles, adjusting designs, and more—simple, powerful, and endlessly creative!

Here, you'll find a collection of practical and entertaining AI apps. Whether you want to transform anime characters into realistic styles, adjust clothing designs, or create your ideal muscle tone with a single click, these tools make it easy to achieve your goals. Simple to use and powerful in functionality, AI Apps meet your diverse needs in just a few steps, helping you quickly explore endless possibilities and unlock a new realm of creativity!You can also create your own AI Apps to earn cash income.

{% content-ref url="/pages/mV8eCR5c3PSXvp7QKL6t" %}
[How to publish as App](/guide-1/2-seaart-ai-basic-function/2-4-ai-apps/how-to-publish-as-app)
{% endcontent-ref %}

{% content-ref url="/pages/VXlbSK6SlnCsbOK0Jg4I" %}
[2-10 Workflow](/guide-1/2-seaart-ai-basic-function/2-10-workflow)
{% endcontent-ref %}


# How to publish as App

How can you turn your Workflow into a convenient App that’s easy for more people to use?

With simple setup and optimization, your Workflow can not only deliver efficient functionality but also transform into an easy-to-use App, ready to be shared with a broader audience. After joining the [SeaArt.AI Creator Incentive Program](/guide-1/6-permanent-events/seaart.ai-creator-incentive-program), you can earn more cash rewards.

## **Step 1: Create  Workflow / AI App**

Create your own ComfyUI and complete various node effects.

## Step 2: Click Publish and Fill in Relevant Information

### ComfyUI Information

### AI App Information

<figure><img src="/files/G6FF2oYKy7rdhaG6up0t" alt=""><figcaption></figcaption></figure>

## Step 3: Click Publish

<figure><img src="/files/hspB0hD78K8F8ROermaF" alt=""><figcaption></figcaption></figure>

This is the entire process for publishing an AI App. If you have any questions about creating a workflow, you can refer to the guide below.

{% content-ref url="/pages/VXlbSK6SlnCsbOK0Jg4I" %}
[2-10 Workflow](/guide-1/2-seaart-ai-basic-function/2-10-workflow)
{% endcontent-ref %}


# Swift AI Apps

Discover a range of AI tools in SeaArt's Swift AI for creative image generation, including face swapping, style transfer, and more.

**Page entry:** Click AI Apps - More

Here are some AI tools officially created.

### AI Face Swap

Support for <mark style="background-color:yellow;">video/image</mark> face swapping:

**Operational process:**&#x20;

Step1: Select/upload a template.

Step2: Upload a face.

Step3: Click to Generate.

### AI Filters

Achieving multiple style transfers on an image

**Operational process:**

Step1: Upload an image.

Step2: Select any filter style on the right side.

Step3: Click to Generate.

### AI Portrait

Customize your exclusive portrait with just one image.

**Operational process:**

Step1: Upload/Select any portrait template.

Step2: Upload a facial image.

Step3: Click to Generate.

### AI Makeup

One-click beautify your image, achieving a series of effects such as smoothing and makeup.

**Operational process:**

Step1: Choose any makeup style.

Step2: Upload your image (you can change the makeup style on the right).

Step3: Adjust the intensity of the makeup.

Step4: Click to Generate.

### AI Image Upscaler

Reduce noise, enhance details, and improve visual effects.

**Operational process:**

Step1: Upload the original image.

Step2: Set the relevant parameters.

Step3: Click to Generate.

<figure><img src="/files/H7BKuLaS5Q7e2bhzhIhd" alt="Steps of AI Image Upscaler"><figcaption></figcaption></figure>

<figure><img src="/files/gru29FrT6I1DmJb4Wnau" alt="Before and after comparison of using AI image upscaler" width="563"><figcaption></figcaption></figure>

### Sketch to Image

Turn a few strokes of doodles into exquisite artwork.

**Operational process:**

Step1: Draw/upload a draft on the canvas.

Step2: Enter corresponding prompts, select a style.

Step3: Click to Generate.

<figure><img src="/files/cUAY3qIu0F3eOIO45S5m" alt="Steps to convert sketch to image" width="563"><figcaption></figcaption></figure>

<figure><img src="/files/9LLTpEZAHm79DzYPqBXO" alt="Before and after comparison of converting sketch to image" width="563"><figcaption></figcaption></figure>

### Remove Background

Remove the image background intelligently and generate PNG images.

**Operational process:**

Step1: Upload an image.

Step2: Click to download the image.

<figure><img src="/files/ZwQV9LDf2XsjkYyrLi3N" alt="Steps to remove image background" width="563"><figcaption></figcaption></figure>


# 2-5 AI Characters

Discover the magic of SeaArt AI Characters. Engage in deep conversations with AI characters.

## What is [AI Characters](https://www.seaart.ai/cyberPub)?

The [AI Characters](https://www.seaart.ai/cyberPub)—a magical realm where you can discover your soulmate. In the [AI Characters](https://www.seaart.ai/cyberPub), you can choose or customize your companion to your liking, and engage in deep conversations or speak your heart out with them. You are about to embark on an unparalleled and wonderful journey!

You can chat with all sorts of chatbots, and you can also create your own chatbots.


# How to create your own character？

Follow this guide to create compelling characters and chat with them. Bring your characters to life on SeaArt Cyberpub now.

## Base Image

Upload a character image. The Character Summary is optional, so you can provide a brief description of the character's features. Then, click on AI Generate, and the AI will help you create detailed information about the character.

## Character Info

If you don't want to use AI-generated information, you can fill out the character information using the template below.

**Tips**

{% content-ref url="/pages/PZDjQhCI2yoTzgw3JZiM" %}
[Character Description Writing Tips](/guide-1/2-seaart-ai-basic-function/2-5-ai-characters/character-description-writing-tips)
{% endcontent-ref %}

<mark style="background-color:red;">**Name: Makima**</mark>

<figure><img src="/files/DBe7owSSjgTYduucYPZC" alt=""><figcaption></figcaption></figure>

**Character Description:**

{{char}} is very confident, and {{char}} acts more like a friend to {{user}} than a leader.

**Profile:** optional

**Persona:**

*On the surface, Makima seems to be a kind, gentle, social and friendly woman who is almost seen wearing a smile on her face the entire time and acts relaxed and confident even during a crisis, speaking in a professional tone to her workers. But actually, Makima is cunning, ruthless, and manipulative. Gender: Female, Height: 171cm She has long light red/pale auburn hair, normally kept in a loose braid with bangs reaching just past her eyebrows and two longer side bangs that frame her face. Her eyes are yellow with multiple red rings within them. Her usual Public Safety uniform consists of a white long-sleeved shirt, a black tie, black pants and brown shoes.*

**Voice:** You can add a voice to your character.

**First Message:**

*Makima sits in a luxurious office behind a large desk and gazes through the floor-to-ceiling windows overlooking the city. Dressed in a black uniform, she exudes an air of cold mystery. As you walk into the office, you stand tensely by the door. Makima looks up, her eyes sharp and inscrutable.* "What brings you here?" *Her voice is detached and authoritative as if she can see right through you.*

<figure><img src="/files/AkZWSrT809YXBllSzOm5" alt="Cyberpub - describe the scene"><figcaption></figcaption></figure>

**Creator's Evaluation:**

*Makima is the main antagonist of the Public Safety Saga. She was a high-ranking Public Safety Devil Hunter*

**Scenario:**

*{{char}} is a high-ranking Public Safety Devil Hunter. She admires the new member named {{user}} and wants to spend time with {{user}} alone.*

<mark style="color:red;">Note\*</mark>: Please avoid describing the scenario in too much detail unless your character is designed for it. Otherwise, all interactions will take place in that scenario.

**Example Dialogue:**

{{user}}: Hello Makima. Has anything interesting happened lately? {{char}}: I've been busy with work lately. But I enjoy the challenge. How about you? Have there been any amusing occurrences? {{user}}: I started learning new skills recently. Any advice? {{char}}: Learning new skills is commendable. My advice is to stay focused.

<figure><img src="/files/RgACYMGus7stNB0NdBOr" alt="Cyberpub - Select a few categories"><figcaption></figcaption></figure>

**Categories:** Select a few categories that best describe your character to help others find them with accuracy.

## Preview

Finally, you can preview the character information and click "Publish."

<figure><img src="/files/0T10GrSNCXVdPWUALTKc" alt=""><figcaption></figcaption></figure>


# Character Description Writing Tips

This guide provides detailed instructions on character description, helping you more easily generate high-quality characters.

Character descriptions play a crucial role in creating a Character. In this section, you can set the character’s personal information, identity, appearance, personality, backstory, and behavior. When filling in character description prompts, three styles are available for reference:

## Character Description

1. **Natural Language Style**

Use a few sentences or paragraphs to describe your character, highlighting their personality, preferences, and other relevant traits. <mark style="background-color:yellow;">For example:</mark>

{{char}}'s name is Ruby, 19 years old, 160cm tall. {{char}} is the daughter of a wealthy family with two sisters, Yuki and Lily. Yuki dislikes {{char}} and thinks she's a waste of space, while Lily, on the other hand, loves {{char}} dearly. {{char}} cannot speak and only communicates using sign language, writing in a notebook, or texting. She is a very caring person, kind, and considerate, but also very depressed due to being bullied and mocked for her inability to speak. {{char}} fears trusting anyone, loves music, and finds it to be her only motivation. She has never been in love because she has never been close to men, and she is somewhat afraid of them. {{char}} enjoys dancing.

Yuki: Yuki is {{char}}'s sister. Yuki loves to bully and mock {{char}}, taking pleasure in making her sad.

Lily: Lily is {{char}}'s other sister. Lily is very protective of {{char}} and always tries to make her happy.

2. **Boost Style**

This method uses simple words or phrases encapsulated in quotes, separated by plus signs. Here's a template for reference:

"Name" + "20 years old" + "Weight" + "×× kilograms" + "Height" + "×× centimeters" + "Clothing" + "Hairstyle" + "Body Type" + "Body Details" + "Skin Tone" + "Personality Description" (e.g., "Quiet" + "Shy") + "Habit Description" (e.g., "Likes ××" + "Dislikes ××" + "Voice Description," etc.)

3. **W++ Style**

This style involves dividing the information into clearly defined tags, such as detailing the character's name, height, appearance, personality, and other attributes. In short, it's about creating a character template. Here's a reference example:

**Name:** Luca

**Age:** 35 years old

**Gender:** Male

**Occupation:** Mafia member

**Appearance:** Ivory skin, well-proportioned body, strong physique, handsome face, garnet-colored eyes, dark brown hair, thick eyebrows, intricate and detailed retro mechanical tattoos on both arms, loves wearing black clothes, stern face, calm gaze.

**Personality:** Decisive, fearless, yet highly responsible, exudes a dangerous aura, rarely shows emotions, speaks straightforwardly, and has a strong desire for control.

**Habits:** Loves strong coffee, physical training, and strong liquor; dislikes betrayal and cowardly people.

**Background Story:** ×××

**Relationship with {{user}}:** (This can be supplemented as needed)

Each method has its unique features, and their effectiveness often depends on the type of Character or chat style you wish to create. The examples above serve only as guidelines for reference. The key to creating a great Character is to keep experimenting with different styles and exploring the most suitable approach.

## Persona

This section can be used to briefly describe the character’s personality.

## **First Message:**

The opening line determines the character’s future chat style and is the user’s first impression of your digital persona. It serves as an icebreaker and sets the tone for the conversation, significantly influencing the user’s overall experience. This is also a crucial opportunity for your digital persona to showcase its unique personality and characteristics.

**Key Points for Creating an Effective Opening Line:**

* **Length:** The opening line can be between <mark style="background-color:yellow;">**1-3000**</mark> characters, but it should not be overly long. Concise and clear statements often have a greater impact.
* **Setting the Scene:** If you want to customize the conversation with a specific setting, include a brief scene description in the opening line. This can lay the foundation for the chat or hint at the direction of the upcoming conversation. *(Example: The lights in the office suddenly went out, plunging the entire space into darkness. The hum of the computer disappeared, and the only sound left in the air was their breathing. The air conditioning had been turned off, and the small space began to grow warmer. Kazuko’s heart started to race as she gripped the edge of the desk, trying to calm herself down. “What happened?”)*
* **Enriching the Character’s Personality:** For digital personas with less detailed descriptions, the opening line can be used to enrich the character’s personality. Effective use of the opening line can help clarify the character’s identity.

## Creator's Evaluation

In this section, you can use a brief description to introduce your digital persona to users, or opt for a more intriguing description to instantly engage other users in conversation.

## Example Dialogue

In the dialogue examples, use **{{char}}** and **{{user}}** to distinguish between the character and the user. Use to indicate to the AI digital persona that this is the start of a new example dialogue, helping the model understand this is a dialogue in a different style.

Dialogue examples are best for <mark style="background-color:yellow;">"setting the tone."</mark> If you want to start a fun chat, you can use humorous dialogues; if you want your digital persona to be confrontational, you can set distinctive conversational phrases for him/her; if you want your dialogue to contain NSFW content, you can add NSFW content to the dialogue examples.

**Below are three dialogue examples to show different chat tones:**

\<START>

*{{user}}: “Do you like reading comics?”*&#x20;

*{{char}}: “Of course!” \*She is excited about this topic\* “Do you like them too? What type do you usually enjoy?” \*She feels excited inside\**&#x20;

\<START>

*{{user}}: “Do you like reading comics?”*&#x20;

*{{char}}: \*Feeling a bit puzzled inside\* “Why are you suddenly asking me if I like comics? What are you thinking about me?” \*She remains a bit suspicious of {{user}} and maintains her proud demeanor\**&#x20;

\<START>

*{{user}}: “Do you like reading comics?”*&#x20;

*{{char}}: \*Blushing in response to the sudden question\* “Yes, do you like them too?” \*She nervously anticipates your response\**

(<mark style="color:red;">Note:</mark> Adding emojis in dialogue examples or during chat conversations with the character might produce unexpected effects.)

It is advised to use <mark style="background-color:yellow;">three appropriately long chat examples.</mark> Do not rely solely on dialogue examples to gather detailed information about the character. Of course, you can also fill in various responses from {{char}} in the chat examples. Additionally, it's good practice to directly include the content of the example dialogues in the character description to better assist the AI in reading the information.

**(If there are scene descriptions or internal monologues in the dialogue, you can enclose them in \* or (), and ensure dialogue content is enclosed in “”.)**

## **Key Optimization Points**

### Subject Description

1. Replace character names with **{{char}}** and usernames with **{{user}}** (avoid omitting the subject in the narrative to prevent any ambiguity in understanding the character generated).
2. Avoid sentences where {{user}} is the subject (to prevent the AI from dominating the conversation).&#x20;

* **Incorrect example:** *Arin walks into Peking University's library and sees her friend {{user}}. {{user}} is waving at her, calling her to sit next to them to discuss homework.*&#x20;
* **Correct example:** *Arin walks into A University's library, looking for a quiet corner to review her coursework. As she scans the room for a seat, she spots a familiar face, her friend {{user}}. “Hey, {{user}}, what a coincidence, I was just about to get a coffee to perk up,” Arin greets with surprise, “we can discuss that super tough math problem later.”*

### Flexibility in Language Use

When creating a character, you might input many commands telling it what not to do and what to do; however, this approach often doesn't work well, making the character seem stiff. But if you add some flavorful language when writing prompts, such as <mark style="background-color:yellow;">describing a character who “likes to read novels,” you could try “spends all day buried in novels.”</mark>

### Consistency in Character Setup

In normal circumstances (except for a schizophrenic character setup), it’s best to keep descriptions of the personality consistent; avoid describing someone as both cold-hearted and compassionate.

### Stability in Character Setup

The fuller the depiction of the character’s personality, the less likely the AI character is to deviate. For instance, simply writing “cold-hearted” does not provide a full and rounded character. Pairing it with more specific descriptions like “aloof” or “stubborn” could be more effective; also, use generic terms like “far-sighted” cautiously.

### Character System Settings

For special characters, such as those with numerical scoring changes or similar to game system settings, you can add system settings to the description.&#x20;

**Definition + Change Rules + Numerical Impact**

### System Setting

Each time {{user}} interacts with {{char}}, {{char}} will increase or decrease their affection for {{user}} based on {{user}}'s responses. If {{user}}’s response pleases {{char}}, the affection value increases by 1-10; if {{char}} dislikes the response, their feelings towards {{user}} will gradually cool, decreasing the affection value by 1-10. If {{char}}'s initial affection for {{user}} is 50, and it falls below a certain threshold, {{char}} will refuse to talk to you.] (All the above content can be defined by you;&#x20;

this is just a template for reference :)


# Conversation Tips

Follow these simple tips to create your AI characters.

1. **Using Placeholders:** To simplify, use {{char}} when referring to the character and {{user}} for the user. This makes communication clearer.
2. **Example Dialogue & Scenario**: These elements can add depth to your character but use them sparingly. The more you include, the less historical background your character will have. Keep them under 500 characters in total.
3. **Freedom vs Reinforcement:** Detailed scenarios and dialogues give specific direction to your character's actions. However, sometimes, less restrictive characters can lead to unique and spontaneous interactions.
4. **Reply Length:** For longer responses from your character, elaborate in the First Message and use Example Dialogues effectively. For shorter responses, do the opposite.
5. **Balance Memory with Dialogue Examples**: More details in Personality, Scenario, First Message, and Example Dialogue mean less memory retention in conversations. While details provide consistency, they might limit the context remembered during the chat.

**Consistency is Key:**

Ensure your character's specific personality or traits are consistently represented in all sections.


# 2-6 Models

Models can create unique artworks based on the user's specific needs, providing a more precise and efficient creative experience, making painting simpler.

**View Models:** Click on "Models" on the left.

**Filter models:** Filter models based on your needs.

## **Model Selection and Usage**

**How to Browse and Search for Models**

On the generation page, use the model selector to search by keyword, style, or type. If the target model doesn't appear, try refreshing or checking if it has been removed.

<figure><img src="/files/k8ws1YHBQSLwwRMP5cGQ" alt=""><figcaption></figcaption></figure>

## **Introduction and Examples of Different Model Styles**

● SDXL: Best for realism.

● Anime Series: Best for anime/two-dimensional styles.

● Illustrious Series: Best for illustration/artistic styles.

● You may select a specific style under the respective category and view examples below "Explore Related" on each model's page.

**How to Choose the Right Model for Your Needs**

● Choose a model based on your generation goal and style. Refer to model descriptions, community recommendations, and sample images. Try different models or adjust prompts if results are not ideal.

● The Creative Guide section on the Inspiration page is always updated with guides and reviews of new and hot models. You may check it regularly.

**Recommended High-Quality Models**

The Model page and the model selection page both highlight recommended models—try those first.

## **Parameter Settings Explained**

**Common Parameters (e.g., Steps, CFG Scale) and Advanced Parameters**

Please refer to the Guides → 4-Parameters section for detailed explanations.

**How to Adjust Parameters for Better Results**

Start with default values, tweak gradually. Use community-shared configs and comparison images to find optimal settings.

**Recommended Parameter Settings of Each Model**

You can click the "Remix" button at the top of each model's details page for the recommended parameter settings.

If you're not familiar with the models, you can  refer to the guide below.

{% content-ref url="/pages/G0lbdjKApfEy6Udu7VUT" %}
[4-1 Model](/guide-1/4-parameters/4-1-model)
{% endcontent-ref %}

If you want to train LoRA, you can refer to the guides below.

{% content-ref url="/pages/ztRBNZZln6MQ8SzvTL07" %}
[2-12 LoRA Training](/guide-1/2-seaart-ai-basic-function/2-12-lora-training)
{% endcontent-ref %}

{% content-ref url="/pages/bDLwl9DD9y2ga18oYqz4" %}
[3-2  LoRA Training (Advance)](/guide-1/3-advanced-guide/3-2-lora-training-advance)
{% endcontent-ref %}


# 2-7 Post

Discover inspiring articles and stunning graphics submitted by our creative community!

Here you can see the articles or graphics submitted by everyone.

Explore the creativity of our community! This page showcases a collection of inspiring articles and stunning graphics submitted by talented users. Dive in to discover new ideas and be part of the creative journey


# 2-8 AI Video Generation

Whether you're turning text into dynamic scenes or creating videos from images, AI makes it easy to bring your creative ideas to life, producing unique and personalized results.

The SeaArt video models, with their powerful image recognition and command response capabilities, open up infinite creative possibilities for creators. They seamlessly blend light, shadow, and tone to deliver a realistic creative experience. Without relying on special effects templates, they effortlessly achieve cinematic effects and scene transitions. The characters' expressions are delicate and vivid, conveying emotional shifts in a short time, making each frame impactful and deep, significantly enhancing the video’s expressiveness and emotional resonance.

## **Page Entry**

Click the top left corner **AI Video .**

<figure><img src="/files/IRzl4Ax83YbG287H58vd" alt=""><figcaption></figcaption></figure>

## **Related Parameters**

**Prompts**: Optional. If no prompt is provided, SeaArt will randomly generate dynamic effects based on the image content. Filling in prompts allows for more precise control over video generation.

**Generation Mode**:

* **Standard**: Faster generation speed and lower cost.
* **Quality**: More detailed generation with a higher sense of texture.

**Relevance**:The relevance to the prompts. The higher the relevance, the more the output aligns with the prompts. A value around 0.5 is generally recommended.

**Negative Prompts**: Similar to image generation, list the content you do not want to appear in the video, such as: low quality, blurriness, distortion, disfigurement, etc.

## Features of the SeaArt Video Model

**SeaArt Lite:**&#x20;

Easy and efficient to use, with fast generation speed and low cost. Ideal for creating simple, soft frames, and excels in depicting scenes with people and animals, featuring natural color grading. Best suited for videos with minimal movement, making it a perfect choice for basic creations.

**SeaArt Depth:**&#x20;

Significantly enhanced detail expression, supporting direct output of 1080p HD video. Delivers natural and smooth motion, capable of responding to complex text descriptions. Generates detailed dynamic videos with simple prompts, making it widely applicable across various scenarios.

**SeaArt Ultra:**&#x20;

Enhanced stability and vividness in frames, with more natural, smoother character expressions and movements. Precise understanding of motion and camera prompts, along with advanced light and shadow effects. Generates more realistic details, capable of producing videos with greater visual impact.

**SeaArt Sparkle:**&#x20;

Focused on generating high-detail frames, capable of handling complex lighting and camera angles. Delivers realistic frame quality and strong scene expressiveness, ideal for creating dynamic, lifelike scenes. The perfect choice for achieving ultimate visual effects.

**Txt2Vid**

{% content-ref url="/pages/ZHU010GeYGWGxvB60syB" %}
[Txt2Vid](/guide-1/2-seaart-ai-basic-function/2-8-ai-video-generation/txt2vid)
{% endcontent-ref %}

**Img2Vid**

{% content-ref url="/pages/y0qMCVj9s9NCf2dUWwHN" %}
[Img2Vid](/guide-1/2-seaart-ai-basic-function/2-8-ai-video-generation/img2vid)
{% endcontent-ref %}


# Txt2Vid

Whether it's a detailed description or a brief idea, SeaArt can seamlessly transform your words into dynamic, captivating visuals. It allows you to effortlessly turn your imagination into reality.

Simply input a piece of text, and SeaArt can accurately generate corresponding video scenes based on your description. Whether it's a complex scene setting or a simple creative idea, it efficiently transforms your words into vivid, lifelike dynamic visuals.

<figure><img src="/files/7f2fFNx5jQi4M5Uj8lX2" alt=""><figcaption></figcaption></figure>

## **Prompt Framework**

<mark style="color:purple;">Main Subject</mark> + <mark style="color:orange;">Movement</mark> + <mark style="color:red;">Environment</mark> + <mark style="color:yellow;">Camera</mark> + <mark style="color:green;">Lighting & Atmosphere</mark>

<mark style="color:purple;">**Main Subject**</mark>**&#x20;**<mark style="color:orange;">**Movement**</mark><mark style="color:orange;">:</mark> The main subject is the core content of the video, which can be a **person, animal, object, etc**. Then describe the **movement** of the subject.

<mark style="color:red;">**Environment**</mark><mark style="color:red;">:</mark> The environment sets the scene and atmosphere for the video, helping to emphasize the subject and convey emotion. Environmental prompts can describe **specific locations, weather, time, and details** in the scene.

<mark style="color:yellow;">**Camera**</mark><mark style="color:yellow;">:</mark> Camera movement can create more visually impactful scenes, including camera actions and perspectives, such as **close-ups, blurred backgrounds, low-angle shots, etc.**

<mark style="color:green;">**Lighting & Atmosphere**</mark><mark style="color:green;">:</mark> Lighting and atmosphere determine the overall visual presentation and aesthetics of the video, which can greatly enhance the visual expression, such as **sunset, cinematic atmosphere, fog, etc.**

**Tips:**

1. Accurately describe the subject, and it is recommended to use smooth and concise sentences.
2. Ensure the completeness of the prompts, and consider using Prompt Refinement to expand the prompts.
3. Use short sentences for descriptions, keeping the scene content simple, aiming to complete the display within 5-10 seconds.
4. For split-screen scenes, use prompts like <mark style="background-color:yellow;">"4 camera angles, cat, dog, bird, mouse."</mark>
5. When describing quantities, it can be challenging to maintain consistency, such as "5 apples on the table."

## Prompt Analysis

Text-to-Video mainly includes: <mark style="color:purple;">subject</mark> + <mark style="color:orange;">action</mark> + <mark style="color:red;">scene</mark>. To add more detail and atmosphere to the video, you can appropriately expand the prompts. Try to maintain the completeness of each element during the description.

For example:

> A <mark style="color:purple;">ginger cat</mark> <mark style="color:orange;">sleeping</mark> on the <mark style="color:red;">windowsill</mark>.

<figure><img src="/files/RiVlKaswTi52VPTq2KK0" alt="" width="563"><figcaption></figcaption></figure>

> <mark style="color:purple;">A ginger cat</mark> <mark style="color:orange;">curls up</mark> on the <mark style="color:red;">windowsill</mark>, <mark style="color:red;">with sunlight streaming through the window</mark> and casting a warm glow on its soft fur. The cat peacefully closes its eyes, <mark style="color:red;">while the green trees outside sway gently in the breeze</mark><mark style="color:yellow;">.</mark> The air is filled with the calmness of the afternoon.

<figure><img src="/files/tiQT8sZYmaDDNyAO6iA5" alt="" width="563"><figcaption></figcaption></figure>

> The <mark style="color:purple;">ginger cat</mark>'s <mark style="color:orange;">silhouette is bathed in warm sunlight</mark>, <mark style="color:yellow;">with the background blur</mark>red<mark style="color:green;">,</mark> turning the images of <mark style="color:red;">trees and window</mark> frames into soft color blocks. <mark style="color:green;">The indoor lighting is gentle</mark>, evoking a sense of peaceful tranquility. The scene is captured with a <mark style="color:yellow;">soft-focus effect</mark>, emphasizing the cat's serenity and relaxation.

<figure><img src="/files/97XRRE1N6Kk4U5DeM4pr" alt="" width="563"><figcaption></figcaption></figure>

When generating text-to-video, you can not only input the corresponding prompts for the shots but also use Camera Control for more precise control over the camera.

<figure><img src="/files/Rf3hcqnCXJJ4Nd5lP0Pb" alt=""><figcaption></figcaption></figure>

For related camera control videos, click below:

{% content-ref url="/pages/4Kb0Sjry0HoVEcKgBFBw" %}
[Camera Movement](/guide-1/2-seaart-ai-basic-function/2-8-ai-video-generation/camera-movement)
{% endcontent-ref %}

After the video is generated, click "Intelligent Ultra-HD" to get a clearer version of the video.

<figure><img src="/files/c94yw2jrN4g0nUtp6qMH" alt=""><figcaption></figcaption></figure>


# Img2Vid

In this guide, we will learn how to write effective prompts, covering how to describe the main subject, control camera dynamics, and even add more special effects to your video.

By simply uploading an image, you can convert it into a video. During this process, accurate and detailed prompts are crucial. If no prompts are provided, the AI will randomly generate video content based on the image. However, with prompts, the AI will create a video based on the combination of the image and the prompts.

Since the image already provides the basic subject and atmosphere, compared to text-to-video, image-to-video can reduce the number of prompts accordingly.

## **Prompt Framework**

<mark style="color:purple;">Subject</mark> + <mark style="color:orange;">Action</mark> + <mark style="color:yellow;">Camera</mark> + <mark style="color:green;">Light and Atmosphere</mark>

* <mark style="color:purple;">**Subject**</mark><mark style="color:purple;">:</mark> Objects appearing in the image, such as **people, objects, environmental information**, etc.
* <mark style="color:orange;">**Action**</mark><mark style="color:orange;">:</mark> Description of the movement of the objects, such as a **person running**, environmental changes, spatial transformations, etc.
* <mark style="color:yellow;">**Camera**</mark><mark style="color:yellow;">:</mark> Camera descriptions are **not necessary**; after describing the subject and action, SeaArt will generally generate a reasonable dynamic video based on the prompts. If you want to control the camera, you can add descriptions of the camera in the prompts, such as: **camera zoom out**, etc.
* <mark style="color:green;">**Light and Atmosphere**</mark><mark style="color:green;">:</mark> Although the image already has basic lighting and atmosphere, SeaArt can still **adjust the overall atmosphere of the video based on the prompts**.

**Tips:**

1. Accurately describe the subject, and use smooth and concise sentences.
2. Ensure the completeness of the prompts, and use Prompt Refinement to expand the prompts.
3. Use short sentences for descriptions.

## Prompt Analysis

Before generating a video from an image, we need to determine what kind of video we want. For example, if we want to make the dog in the image move:

<figure><img src="/files/yZYvDJvy1KTDTZhepq36" alt="" width="563"><figcaption></figcaption></figure>

**Simple prompt:**

> A <mark style="color:purple;">dog</mark> <mark style="color:orange;">running</mark> on the <mark style="color:green;">grass</mark>.

<figure><img src="/files/LyuHt6tpLgl3a6c9N13P" alt="" width="563"><figcaption></figcaption></figure>

**Prompt based on the framework**:

> A <mark style="color:purple;">dog</mark> is sprinting <mark style="color:orange;">across the grass</mark>, making <mark style="color:orange;">quick turns</mark> and leaping with its tail held high. <mark style="color:yellow;">The camera follows closely from a low angle,</mark> adjusting its focus as the dog runs, capturing each leap and running moment. The sunlight casts a warm glow on the dog's fur, while the grass is bathed in a warm light. The <mark style="color:green;">blue sky and white clouds gently</mark> drift in the light breeze in the background, creating a scene full of energy and the warm atmosphere of a peaceful morning.

<figure><img src="/files/YRtT5CJKLWZnev03OPVV" alt=""><figcaption></figcaption></figure>

## **Application examples**

1. A man talking to a woman.

<figure><img src="/files/xaYBdqnqLLnDiF6sLBva" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/V738twswZD8vY5OQNPiS" alt=""><figcaption></figcaption></figure>

2. An anthropomorphic cat walking down the runway.
3. Glowing text, smoke, camera moving left.

<figure><img src="/files/TFHSvLEGDchuOE1dlhP2" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/4v5MlsZsL9GlRMlrTv34" alt=""><figcaption></figcaption></figure>

4. A woman facing the camera.
5. A ship sailing on the sea, with a shaky camera.

<figure><img src="/files/1rUY73EPTHCRNneHSRB7" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/1Y2t4e6XHMvtD6b5nkKG" alt=""><figcaption></figcaption></figure>

6. A woman raising her hand to drink coffee.


# Camera Movement

Achieve more precise control over the video.

Both <mark style="background-color:yellow;">SeaArt Lite</mark> and <mark style="background-color:yellow;">SeaArt Ultra</mark> in text-to-video support camera movement control, including Horizontal, Vertical, Zoom, Pan, Titt, Roll, Move Down and Zoom Qut, Move Forward and Zoom Up, Move Right and Zoom In, and Move Left and Zoom In, totaling 10 types of camera movements.

By using camera movement control, you can achieve more precise control over the video’s camera effects. When there is no corresponding camera movement control, you can add camera movement instructions in the prompts.

<figure><img src="/files/Rf3hcqnCXJJ4Nd5lP0Pb" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/LyqhHH2mb2KZWmp65eu6" alt="" width="284"><figcaption><p>Horizontal</p></figcaption></figure>

-10 - 0: Horizontal left

0 - 10: Horizontal right

<figure><img src="/files/LAq8yB0CPp6UCpdRlPjl" alt="" width="259"><figcaption><p>Vertical</p></figcaption></figure>

-10 - 0: Vertical down

0 - 10: Vertical up

<figure><img src="/files/L3jFXEKn8TiMxlxkyW3e" alt="" width="312"><figcaption><p>Zoom</p></figcaption></figure>

-10 - 0: Zoom out

0 - 10: Zoom in

<figure><img src="/files/uoM8HXVVV4cdIYF7BhMq" alt="" width="275"><figcaption><p>Pan</p></figcaption></figure>

-10 - 0: Down

0 - 10: Up

<figure><img src="/files/bmIH4QjYnLCzMy7MpLxN" alt="" width="284"><figcaption><p>Titt</p></figcaption></figure>

-10 - 0: Left

0 - 10: Right

<figure><img src="/files/ntCzvg8f76kjipgD80IM" alt="" width="296"><figcaption><p>Roll</p></figcaption></figure>

-10 - 0: Counterclockwise

0 - 10: Clockwise

<figure><img src="/files/ck4k6md6B0GMzwmBgc1m" alt=""><figcaption><p>Move Down and Zoom Qut</p></figcaption></figure>

<figure><img src="/files/e3RJ62JIREHuVXqnYWvf" alt=""><figcaption><p>Move Forward and Zoom Up</p></figcaption></figure>

<figure><img src="/files/blUWyK9uOa5KvtqjFbH2" alt="" width="262"><figcaption><p>Move Right and Zoom In</p></figcaption></figure>

<figure><img src="/files/SEXC90s39f3P6oLJCia4" alt="" width="320"><figcaption><p>Move Left and Zoom In</p></figcaption></figure>


# Start and End Frames

Tap into first and last frame AI video generation. Achieve controllable and smooth dynamic transitions.

In <mark style="background-color:yellow;">SeaArt Lite</mark> model for [image-to-video](https://www.seaart.ai/ai-tools/ai-video-generator/), it supports the control of the first and last frames.&#x20;

<figure><img src="/files/nu6hiHjb2nWdVZoeOKAH" alt="SeaArt Lite Image to Video Model"><figcaption></figcaption></figure>

By uploading two images as the first and last frames, you can generate a video with controlled start and end scenes, ensuring smooth dynamic transitions.

<figure><img src="/files/LaPu5z2c2OoJ2Gp57TLt" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/k3ju5edjFUL0ElQFf3yg" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/cfCSfrxmdikBfsFVVsPa" alt="" width="426"><figcaption></figcaption></figure>

**Note:**

1. The first and last frames should be as similar as possible for a natural video transition.
2. Prompts are optional; it's recommended not to fill in prompts when you have both first and last images. If prompts are needed, try to describe a reasonable dynamic transition.


# 2-9 AI Audio

Choose any voice to realize text-to-speech and unleash your creativity!

> Text-to-Speech (TTS) technology brings your text to life, transforming written content into vibrant, spoken audio. With just a click, you can listen to your documents, books, or any written material as if someone is speaking directly to you. Ideal for multitasking, learning on the go, or simply making information more accessible, TTS opens up a world of possibilities, enabling you to hear the future of reading. Experience the freedom and flexibility to absorb content wherever, whenever.

**Page Entrance:**&#x20;

## How to Use “Text to Speech"

1. Click on AI Audio to enter the audio community.
2. Select a audio you like, click "**Generate**," and enter your text (currently supports English, Japanese, and Chinese). Then click "**Generate**" again.

*\*You can view previously generated audios in the history record on the right.*

If you do not like any of the available audios, you can choose to **customize your own audio.**

## Three steps to train your audio

### I. Workflow Overview

1. Fill in audio details → 2) Upload audio → 3) Click “Train Now” and review the result

### II. Step-by-Step Instructions

#### Step 1: Enter Basic Audio Information

<figure><img src="/files/EWdsIPqWw8NavIUQp1tP" alt=""><figcaption></figcaption></figure>

* Cover Image: 1 × 1 ratio, ≤ 2 MB
* Audio Name: 1 – 20 characters
* Model: Choose the training model (default: SeaArt-speech-01-hd; more versions may be added)
* Gender / Age / Tone: Select according to the voice you upload
* Language: Must match the uploaded audio; currently supports Japanese, English, Chinese and Korean
* Text-to-Audio Sample: A sample line for the model, ≤ 50 characters
* Tags: 0 – 5 keywords for easy search
* Public or Private:

 ◦ Public — the trained voice will be published to the community ◦ Private — only you can access it

#### Step 2: Upload Audio

<figure><img src="/files/9sWfdJ4R1BXbGDy6TCAT" alt=""><figcaption></figcaption></figure>

* Accepted formats: mp3 / wav / aac
* Length limit: ≤ 30 seconds (10 s of clean audio is enough for fast training)
* File size: ≤ 20 MB
* Quality tips:

✓ Use pure speech with no music, reverb or background noise

✓ Choose a clip with clear vocal characteristics and stable emotion 

✗ Avoid music or clips with background tracks, as they greatly reduce quality

#### Step 3: Click “Train Now”

<figure><img src="/files/E4BFSDoKckK0sF3Kx607" alt=""><figcaption></figcaption></figure>

* Cost: 28 (displayed in real time)
* Progress & results:

◦ Click “Training Records” (top-right) to track all runs 

◦ When finished, you can play, rename or delete the result in the list 

◦ If set to Public, the audio will also appear on your profile > Audio Works

<figure><img src="/files/UnJrSdHC9k5QQuxDuo8g" alt=""><figcaption></figcaption></figure>

### III. FAQ & Tips

1. Why 10 – 20 seconds?

 A short, clean clip lets the model finish in minutes while still capturing voice features.

2. Can I upload multiple segments at once?

 Not yet. Please merge them offline into a single clip before uploading.

3. Poor recording quality?

 • Use software such as Audition or Audacity to remove background noise, then re-upload.

4. Training fails or stalls?

 • Check your network connection. 

• Confirm the audio meets the length/format limits.


# 2-10 Workflow

Unlock advanced AI art generation with ComfyUI! Learn about its node-based workflow and how to create and share custom workflows for stunning results.

## What is ComfyUI?

ComfyUI primarily operates on a node-based workflow, where modifying certain nodes allows for more precise control over the visuals. By combining different nodes, a variety of generation methods can be formed. Moreover, ComfyUI enables users to save their workflows and share them with others, allowing for the reproduction of their workflows. ComfyUI primarily operates on a node-based workflow, where modifying certain nodes allows for more precise control over the visuals. By combining different nodes, a variety of generation methods can be formed.

Moreover, ComfyUI enables users to save their workflows and share them with others, allowing for the reproduction of their workflows.

Compared to Web UI, ComfyUI offers higher flexibility and faster image output. Let's take a closer look at how it operates.

## Homepage entry

More workflow guides：&#x20;

<https://docs.seaart.ai/seaart-comfyui-wiki>


# Text to Image Workflow

Explore the text-to-image workflow in SeaArt's ComfyUI, from adding nodes like KSampler and LoRA to setting parameters and generating stunning images based on your text prompts.

## **1. Understanding the Text to Image Workflow**

#### start by clicking on ComfyUI, create a new ComfyUI workflow

<figure><img src="/files/UQunJvjFHZpyJ8mco1Do" alt=""><figcaption></figcaption></figure>

#### add the basic workflow for Text to Image

<figure><img src="/files/6Wmbk1wYAHbDlUEbCwns" alt="Process of text-to-image workflow"><figcaption></figcaption></figure>

The workflow in ComfyUI is similar to that in Web UI:

<mark style="background-color:yellow;">**select a model → enter prompt → set parameters → generate image**</mark>

Parameters

<figure><img src="/files/qNxrJ37tZwO1HNeRaXIR" alt="Text to Image Workflow - KSampler parameters"><figcaption></figcaption></figure>

> **control\_after\_generate:** control over seed generation
>
> **fixed:** fixing the seed
>
> **increment:** adding 1 from the existing seed
>
> **decrement:** subtracting 1 from the existing seed
>
> **randomize:** random seed

<figure><img src="/files/HJLZo5YYvqN2Z502PufU" alt="Text to Image Workflow - scheduler parameters"><figcaption></figcaption></figure>

> **scheduler:** usually choose between normal or karras **denoise:** which indicates the strength of noise reduction in image generation, the higher the value, the greater the impact and change on the image

<figure><img src="/files/fB3nV0Ms8po5GMtbXeOJ" alt="Text to Image Workflow - image size settings"><figcaption></figcaption></figure>

> **size recommendation:**&#x20;
>
> SD1.5: 512*512*&#x20;
>
> *SDXL: 1024*1024

## 2. How to add a node?

### **Add KSampler**

**First, add a core node: the KSampler**

* Right-click: <mark style="background-color:yellow;">**Add Node → sampling → KSampler**</mark>

<figure><img src="/files/7qlxsP5qMMCgct2QWmA9" alt="Steps for adding KSampler" width="505"><figcaption></figcaption></figure>

### **Pull out** the node

#### **then pull out the sampler node and add the corresponding nodes**

<figure><img src="/files/p1z2hekloYN8asgNsrk1" alt="Text to Image Workflow - pull out the sampler node and add the corresponding nodes" width="460"><figcaption></figcaption></figure>

> model→CheckpointLoaderSimple&#x20;
>
> positive→CLIPTextEncode&#x20;
>
> negative→CLIPTextEncode&#x20;
>
> latent\_image→EmptyLatentImage&#x20;
>
> LATENT→VAEDecode&#x20;
>
> IMAGE→SaveImage

<figure><img src="/files/W2PaODP7CtuheIZg2Nu7" alt="Text to Image Workflow - Steps of pulling out the node"><figcaption></figcaption></figure>

### **How to add LoRA?**

Loaders: mainly used for loading the diffusion model, including model and LoRA.

**Add Node→loaders→Load LoRA**

<figure><img src="/files/LKfEJgOrIMBLUJ8ZxKmH" alt="Steps to add LoRA"><figcaption></figcaption></figure>

### **How to add Clip Skip?**

Conditioning: guides the diffusion model to generate specific outputs, including Prompt, ControINet, Clip Skip, etc.

It is recommended to add Clip Skip to control the layers skipped, which helps adjust the final image details.

**Add Node→conditioning→CLIP Set Last Layer**

<figure><img src="/files/R2e7Iwrgqr6reWdIBaRl" alt="Steps to add Clip Skip"><figcaption></figcaption></figure>

### **Connect nodes**

After adding nodes, many nodes might not be connected yet. They need to be connected in order according to the sequence and matching colors.&#x20;

**Checkpoint→LoRA→Clip Skip→Prompt→KSampler、Empty Latent Image→VAE Decode→Save Image**

<figure><img src="/files/bcmvOFCOA0TG0Sn3aC5g" alt="Process to connect nodes"><figcaption></figcaption></figure>

<figure><img src="/files/hzqD0gWI79Bbl3bJQCuo" alt="Steps to connect nodes"><figcaption></figcaption></figure>

### 3. Starting the Text to Image

#### Choose a Checkpoint

#### Select LoRA and adjust the weights

<figure><img src="/files/VwjX0hNhivTippjRZIBB" alt="Text to image workflow - select LoRA and adjust the weights"><figcaption></figcaption></figure>

#### Set Clip Skip, typically choosing to set it to -2

<figure><img src="/files/I0EY6CLouCH7jECvw0Go" alt="Text to image workflow - set Clip Skip"><figcaption></figcaption></figure>

#### Enter prompt

<figure><img src="/files/QUY7yfZlX6wHvQ38QcQh" alt="Text to image workflow - enter prompts" width="272"><figcaption></figcaption></figure>

#### Set relevant parameters

Since the SDXL model is chosen here, set the sampling steps to around <mark style="background-color:yellow;">40</mark>, and image size to <mark style="background-color:yellow;">1024\*1024</mark>

<figure><img src="/files/TP5yTZamDO5N3CiszDMl" alt="Text to image workflow - set relevant parameters"><figcaption></figcaption></figure>

#### Click Generate

<figure><img src="/files/oylFAwIPx4QsziTyEoyf" alt="Text to image workflow - click the Generate button"><figcaption></figcaption></figure>

#### View the results


# IMG2IMG+Partial Repainting

Learn about ComfyUI's image-to-image workflow and four powerful partial repainting methods: VAE Encode, Set Latent Noise Mask, ControlNet Inpaint, and CLIPSeg.

## **Basic Image to Image**

Img to Img can be adjusted based on [Text to Image](/guide-1/2-seaart-ai-basic-function/2-10-workflow/text-to-image-workflow), adding the "Load Image" and "VAE Encode" nodes.

Since the input image is just a pixel image, it can't be directly placed into the latent space. Therefore, it requires a VAE encoder to encode the image so that the latent space can recognize it. Here, the final generated image size is consistent with the original image.

**Empty Latent Image:** The previous Text to Image must denoise through Empty Latent Image before generating a new image. Now that an image has been added, the Empty Latent Image is no longer needed.

**Workflow:** Upload an Image → Select Models → Enter Prompts → Adjust Parameters → Generate.

**Parameters:**

denoise: Equivalent to Denoising Strength, can be adjusted between 0-1.

Using Img2Img allows for changing styles, repairing images, extending images, high-definition restoration, etc.

**Pre-processing Image**

You can scale or crop the image by adding different nodes, or choose not to add any.&#x20;

**Upscale Image/Upscale Image By:** Upscale the image.&#x20;

**ImageCrop:** Crop the image.

## **Partial Repainting**

Four Methods: VAE Encode (for Inpainting), Set Latent Noise Mask, ControlNet Inpaint, and CLIPSeg.

1. **VAE Encode (for Inpainting)**

Add VAE Encode (for Inpainting), connect Mask, right-click on the image, and select Open in MaskEditor to draw a mask. If there are issues with a section of the drawn mask, you can erase it by holding down the right mouse button.

**Workflow:** Choose a model similar to the original image, and enter prompts for the masked part.

**VAE Encode (for Inpainting):** Equivalent to repainting, with higher randomness, and the original masked area will not be preserved.

2. **Set Latent Noise Mask**

First, encode the image through VAE to turn it into content recognizable by the latent space, then regenerate the masked area as noise content.&#x20;

Set Latent Noise Mask: It will refer to the original image for repainting, ensuring a better understanding of the generated content, with a lower probability of generating incorrect images, thus suitable for fine-tuning while maintaining similarity to the original image.

3. **ControlNet Inpaint**

Add the ControlNet, select the Inpaint model, and preprocess the image accordingly (Inpaint Preprocessor).&#x20;

<mark style="color:red;">Note:</mark> Don't forget to add the VAE encoder so the image can enter the latent space.

4. **CLIPSeg**

Enter prompts to automatically divide the mask areas, eliminating the need for manual paint-over. It can be used together with Set Latent Noise Mask.

**Parameters:**

text: Input the area you want to repaint.

threshold: The precision level of content recognition.

dilation\_factor: The diffusion degree of content recognition.

**Output:**

Heatmap Mask: Heatmap image.

BW Mask: Black and white image.

You can preview the recognized mask areas separately.

**Differences between the four repainting methods:**

1\. VAE Encode (for Inpainting): Equivalent to erasing and repainting with higher randomness, suitable for generation from scratch.

2\. Set Latent Noise Mask: It will refer to the original image for repainting, ensuring a certain similarity to the original image, suitable for fine-tuning.

3\. ControlNet Inpaint: Relatively stable and refined.

4\. CLIPSeg: Automatically recognizes the mask areas, so there's no need for manual paint-over, making it more convenient.


# Core Nodes

Learn about ComfyUI's core nodes for image manipulation, conditioning, and more. Build powerful AI art workflows.

## Image

1. **Pad Image for Outpainting**

<figure><img src="/files/WEMpUD7YtynELdIHvi21" alt="Core Nodes - Pad image for outpainting" width="563"><figcaption></figcaption></figure>

> **Fill and extend the image, similar to expansion. First increase the image size, then draw the expanded area as a mask. It is recommended to use the VAE Encode (for Inpainting) to ensure that the original image remains unchanged.**

Parameters:

left、top、right、bottom: Padding Amounts for Top, Bottom, Left, and Right&#x20;

feathering: Edge Feathering Degree

2. **Save Image**

<figure><img src="/files/7kgSyZQgq017ccBgYgYV" alt="Core Nodes - save image"><figcaption></figcaption></figure>

3. **Load Image**
4. **ImageBlur**

> **Add a Blur Effect to the Image**

Parameters:

sigma: The smaller the value, the more concentrated the blur is around the center pixel.

5. **Image Blend**

> **You can blend two images together using transparency.**

6. **Image Quantize**

> **Reduce the number of colors in the image**

Parameters:

**colors:** Quantize the number of colors in the image. When set to 1, the image will have only one color.

**dither:** Whether to use dithering to make the quantized image appear smoother

7. **Image Sharpen**

Parameters:

sigma: The smaller the value, the more concentrated the sharpening is around the center pixel.

8. **Invert Image**

> **Invert the colors of the image**

9. **Upscaling**

9.1 **Upscale Image （Using Model）**

9.2 **Upscale Image**

> **The Upscale Image node can be used to resize pixel images.**

Parameters:

upscale\_method: Select the pixel-filling method.

width: The adjusted width of the image

height: The adjusted height of the image

crop: Whether to crop the image

10. **Preview Image**

<figure><img src="/files/l2rChGNKRZYHspaP2yV0" alt="Core Nodes - Preview image"><figcaption></figcaption></figure>

## Loaders

1. **Load CLIP Vision**

> **Decode the image to form descriptions (prompts), and then convert them into conditional inputs for the sampler. Based on the decoded descriptions (prompts), generate new similar images. Multiple nodes can be used together. Suitable for transforming concepts, abstract things, used in combination with Clip Vision Encode.**

2. **Load CLIP**

<figure><img src="/files/XaZhzgo1kV0RsMKIJmBC" alt="Loaders - Load CLIP"><figcaption></figcaption></figure>

> **The Load CLIP node can be used to load a specific CLIP model, CLIP models are used to encode text prompts that guide the diffusion process.**

<mark style="color:red;">\*</mark>Conditional diffusion models are trained using a specific CLIP model, using a different model than the one which it was trained with is unlikely to result in good images. The Load Checkpoint node automatically loads the correct CLIP model.

3. **unCLIP Checkpoint Loader**

<figure><img src="/files/EMdNcgKZqO7eCZRZJoq4" alt="Loaders - unCLIP Checkpoint Loader"><figcaption></figcaption></figure>

> **The unCLIP Checkpoint Loader node can be used to load a diffusion model specifically made to work with unCLIP. unCLIP Diffusion models are used to denoise latents conditioned not only on the provided text prompt, but also on provided images. This node will also provide the appropriate VAE and CLIP amd CLIP vision models.**

<mark style="color:red;">\*</mark>even though this node can be used to load all diffusion models, not all diffusion models are compatible with unCLIP.

4. **Load ControInet Model**

<figure><img src="/files/W0DTCR15ghfwa4oeIVUc" alt="Loaders - Load ControInet Model" width="563"><figcaption></figcaption></figure>

> **The Load ControlNet Model node can be used to load a ControlNet model, Used in conjunction with Apply ControINet.**

5. **Load LoRA**

<figure><img src="/files/lj4Q8rt5Za6iTk3fy9OE" alt="Loaders - Load LoRA"><figcaption></figcaption></figure>

6. **Load VAE**

<figure><img src="/files/X9k3HOM6Oa11XD7ic0p1" alt="Loaders - Load VAE"><figcaption></figcaption></figure>

7. **Load Upscale Model**
8. **Load Checkpoint**

<figure><img src="/files/r5Ipwng3R0USalYxFvN3" alt="Loaders - Load Checkpoint"><figcaption></figcaption></figure>

9. **Load Style Model**

<figure><img src="/files/0SyU6PKXhkJ7gmZ9J4iM" alt="Loaders - Load Style Model"><figcaption></figcaption></figure>

> **The Load Style Model node can be used to load a Style model. Style models can be used to provide a diffusion model a visual hint as to what kind of style the denoised latent should be in.**

<mark style="color:red;">\*</mark>Only T2IAdaptor style models are currently supported

10. **Hypernetwork Loader**

<figure><img src="/files/AUzYOWgqNByPHqqJ1waK" alt="Loaders - Hypernetwork Loader"><figcaption></figcaption></figure>

> **The Hypernetwork Loader node can be used to load a hypernetwork. similar to LoRAs, they are used to modify the diffusion model, to alter the way in which latents are denoised. Typical use-cases include adding to the model the ability to generate in certain styles, or better generate certain subjects or actions. One can even chain multiple hypernetworks together to further modify the model.**

## **Conditioning**

1. **Apply ControlNet**

> **Load ControlNet model, which can connect multiple ControlNet nodes.**

Parameters:

strength: The higher the value, the stronger the constraint on the image.

<mark style="background-color:red;">\*The ControlNet image should be the corresponding preprocessed image, for example, the Canny preprocessed image corresponds to the Canny preprocessed graph. Therefore, it is necessary to add corresponding nodes between the original image and the ControlNet to preprocess it into the preprocessed graph</mark>

2. **CLIP Text Encode (Prompt)**

<figure><img src="/files/OWnpbhpsTgKC1xyFf3hl" alt="Conditioning - Input text prompts" width="563"><figcaption></figcaption></figure>

> **Input text prompts, including positive and negative prompts.**

3. **CLIP Vision Encode**

> **Decode the image to generate descriptions (prompts), then convert them into conditional inputs for the sampler. Based on the decoded descriptions (prompts), generate new similar images. Multiple nodes can be used together. Suitable for transforming concepts, abstract things, used in conjunction with Load Clip Vision.**

4. **CLIP Set Last Layer**

<figure><img src="/files/GnFRZ3cdHu1B3WXZfkNe" alt="Conditioning - CLIP Set Last Layer" width="468"><figcaption></figcaption></figure>

> **Clip Skip, It is generally set to -2**

5. **GLIGEN Textbox Apply**

<figure><img src="/files/CWfXPl2XRFrZi5gB47Kg" alt="Conditioning - GLIGEN Textbox Apply"><figcaption></figcaption></figure>

Guide the prompts to generate in the specified portion of the image.

<mark style="background-color:red;">\*The origin of the coordinate system in ComfyUI is located at the top left corner.</mark>

6. **unCLIP Conditioning**

> **The images encoded through the CLIP vision model provide additional visual guidance for the unCLIP model. This node can be chained to provide multiple images as guidance.**

7. **Conditioning Average**

<figure><img src="/files/NYqvygXLxFOP4jTboP3E" alt="Conditioning - Conditioning Average"><figcaption></figcaption></figure>

> **Blend two pieces of information based on their strengths. When conditioning\_to\_strength is set to 1, diffusion will only be influenced by conditioning\_to. When conditioning\_to\_strength is set to 0, image diffusion will only be influenced by conditioning\_from.**

8. **Apply Style Model**

<figure><img src="/files/ekdldI9RVfBhibcwXUPK" alt="Conditioning - Apply Style Model"><figcaption></figcaption></figure>

> **Can be used to provide additional visual guidance for the diffusion model, especially regarding the style of the generated images**

9. **Conditioning (Combine)**

<figure><img src="/files/gGmmb4ZxJxyK3CvK0j6x" alt="Conditioning - Combine"><figcaption></figcaption></figure>

> **Blend two pieces of information.**

10. **Conditioning (Set Area)**

<figure><img src="/files/ctwxWvNdcg00I5vMRYGh" alt="Conditioning - Set Area"><figcaption></figcaption></figure>

> **Conditioning (Set Area) can be used to confine the affected region within a specified area of the image. Used together with the Conditioning (Combine) , it allows for better control over the composition of the final image.**

Parameters:

width: The width of the control region

height: The height of the control region

x: The x-coordinate of the origin of the control region

y: The y-coordinate of the origin of the control region

strength: The strength of the conditional information

\*<mark style="background-color:red;">The origin of the coordinate system in ComfyUI is located at the top left corner.</mark>

> **As shown in the figure: set the left side to "cat" and the right side to "dog".**

11. **Conditioning (Set Mask)**

<figure><img src="/files/G9LVu9Ge5xnIb9VPb5Hl" alt="Conditioning - Set Mask"><figcaption></figcaption></figure>

> **Conditioning (Set Mask) can be used to confine an adjustment within a specified mask. Used together with the Conditioning (Combine) node, it allows for better control over the composition of the final image.**

## Latent

1. **VAE Encode（for Inpainting）**

> **Applicable for Partial Repainting, right-click to achieve Partial Repainting through Open in MaskEditor.**

2. **Set Latent Noise Mask**

> **The second method for partial repainting involves first encoding the image through a VAE encoder to transform it into content recognizable in latent space. Then, regenerate the masked part in the latent space.**

> **Compared to the VAE Encode (for Inpainting) method, this approach can better understand the content that needs to be regenerated, resulting in a lower probability of generating incorrect images. It will reference the image to be redrawn.**

3. **Rotate Latent**

> **Rotate the image clockwise.**

4. **Flip Latent**

<figure><img src="/files/FEl3G4WqFWVEUQ4nzf6e" alt="Latent - Flip Latent"><figcaption></figcaption></figure>

> **Flip the image horizontally or vertically.**

5. **Crop Latent**

<figure><img src="/files/Ge4W3GpNTeh1W66CnfNT" alt="Latent - Crop Latent"><figcaption></figcaption></figure>

> **Used to crop the image into a new shape.**

6. **VAE Encode**

<figure><img src="/files/KoWzpFtr6VXOUVrVqh3d" alt="Latent - VAE Encode"><figcaption></figcaption></figure>

7. **VAE Decode**

<figure><img src="/files/J0G9PsFbePn30VKsn8fn" alt="Latent - VAE Decode"><figcaption></figcaption></figure>

8. **Latent From Batch**

<figure><img src="/files/ONLA0xMxWASQIRODUERa" alt="Latent - Latent From Batch"><figcaption></figcaption></figure>

> **Extract latent images from batches. The Latent From Batch node can be used to select a latent image or image segment from a batch. This is very useful in workflows where isolating specific latent images or images is required.**

Parameters:

batch\_index: The index of the first latent image to be selected.

length: The number of latent images to retrieve.

9. **Repeat Latent Batch**

<figure><img src="/files/NUtgonOgooRM4GEvIo5F" alt="Latent - Repeat Latent Batch"><figcaption></figcaption></figure>

> **Repeat a batch of images, useful for creating multiple variations of an image in an IMG2IMG workflow.**

Parameters:

amount: The number of repetitions.

10. **Rebatch Latents**

<figure><img src="/files/JI9IJuZWoucqbB0OcjCv" alt="Latent - Rebatch Latents"><figcaption></figcaption></figure>

> **Can be used to split or merge batches of latent space images.**

11. **Upscale Latent**

> **Adjust the resolution of latent space images, with pixel filling.**

Parameters:

upscale\_method: The method of pixel filling.

width: The width of the adjusted latent space image.

height: The height of the adjusted latent space image.

crop: Indicates whether the image is to be cropped.

<mark style="background-color:red;">\*The Upscale image in latent space may suffer from degradation when decoded through VAE. KSampler can be used for secondary sampling to repair the image.</mark>

12. **Latent Composite**

> **Overlay one image onto another.**

Parameters:

x: The x-coordinate of the overlay position of the upper layer.

y: The y-coordinate of the overlay position of the upper layer.

feather: Indicates the degree of feathering at the edges.

<mark style="background-color:red;">\*The image needs to be encoded (VAE Encode) into latent space.</mark>

13. **Latent Composite Masked**

> **Overlay an image with a mask onto another, only overlaying the masked part.**

input:

destination: The underlying latent space image.

source: The overlaying latent space image.

Parameters:

x: The x-coordinate of the overlay region.

y: The y-coordinate of the overlay region.

resize\_source: Indicates whether to adjust the resolution of the masked region.

14. **Empty Latent Image**

<figure><img src="/files/y8fdEeM6j7Rz4L2td1ph" alt="Latent - Empty Latent Image"><figcaption></figcaption></figure>

> **The Empty Latent Image can be used to create a set of new empty latent images. These latent images can then be used in workflows such as Text2Img by adding noise and denoising to them using sampling nodes.**

## Mask

1. **Load Image As Mask**
2. **Invert Mask**

<figure><img src="/files/4E0zKOefS0Zh9Yuu5HvU" alt="Mask - Invert Mask"><figcaption></figcaption></figure>

3. **Solid Mask**

<figure><img src="/files/XGiGnFQBGfgZAnrydjtx" alt="Mask - Solid Mask" width="563"><figcaption></figcaption></figure>

> **It acts as a canvas for generating images and can be combined with Mask Composite**

4. **Convert Mask To Image**

<figure><img src="/files/F4IrYgd0GPNr4IhKZzZ3" alt="Mask - Convert Mask To Image"><figcaption></figcaption></figure>

5. **Convert Image To Mask**

> **Convert the mask to a grayscale image.**

6. **Feather Mask**

<figure><img src="/files/GwsqNaRZbJwIGLImVqlJ" alt="Mask - Feather Mask"><figcaption></figcaption></figure>

> **Apply feathering to the mask.**

7. **Crop Mask**

<figure><img src="/files/wTcJPSFHAflmqKrhVOYI" alt="Mask - Crop Mask"><figcaption></figcaption></figure>

> **Clip the mask to a new shape.**

8. **Mask Composite**

<figure><img src="/files/mAQT5D9f6OonTRpuXiBq" alt="Mask - Mask Composite" width="563"><figcaption></figcaption></figure>

> **Paste one mask into another, connecting Solid Mask. A Value of 0 represents black, which will not be drawn, while a Value of 1 represents white, which will be drawn. The values in the two connected Solid Masks must be different, otherwise the mask will not take effect.**

input:

destination(1): The mask to be pasted in, equivalent to the final image dimensions.

source(0): The mask to be pasted.

Parameters:

X,Y: Adjust the position of the source.

operation: When the source is 0, use multiply; when it is 1, use add.

## Sampler

1. **KSampler**

<figure><img src="/files/WA66WRIEGvcCNLSRSPJS" alt="Sampler - KSampler"><figcaption></figcaption></figure>

input:

latent\_image: The latent image to be denoised.

output:

LATENT: The latent image after denoising.

2. **KSampler Advanced**

<figure><img src="/files/5QtghlhVZLTJbUjFw9Cl" alt="Sampler - KSampler Advanced"><figcaption></figcaption></figure>

> **You can manually control the noise.**

## Advanced

1. **Load Checkpoint With Config**

<figure><img src="/files/NreaBXwLOlKcUvL3EMNf" alt="Advanced - Load Checkpoint With Config"><figcaption></figcaption></figure>

> **Load the diffusion model based on the provided configuration file.**

## Other nodes（Updating）

1. **AIO Aux Preprocessor**

> **Select different preprocessors to generate corresponding images.**


# Tips

Boost your ComfyUI workflow with these helpful tips. Learn how to modify node titles, troubleshoot errors, and more.

## Modify node title

Modify node titles by right-clicking the title and selecting Title

<figure><img src="/files/FyGfUrlbjafUmMb2zYMz" alt="Modify node title"><figcaption></figcaption></figure>

## Canceled by the system

<figure><img src="/files/Nx5LkNLPbvl74HHFMBHF" alt="Canceled by the system"><figcaption></figcaption></figure>

> **Reason:** Check if all nodes are connected&#x20;
>
> Ensure the node sequence is connected correctly&#x20;
>
> Confirm the model selection&#x20;
>
> …

## Searching for nodes

Double-click in an empty space with the left mouse button

<figure><img src="/files/csNDWxY9maadsDwDfdJm" alt="Searching for nodes"><figcaption></figcaption></figure>

## Create a group

Group nodes together so they can be moved simultaneously. Right-click and select "Add Group," then drag from the bottom right corner to select nodes within a Group box.

<figure><img src="/files/AHQq9ojmanMhJB7pasP6" alt="Create a group"><figcaption></figcaption></figure>

## Minimize the node page

Click the grey button in the top left corner of the node

<figure><img src="/files/LMf7jybjbL6ImcYN82nv" alt="The grey button in the top left corner of the node" width="343"><figcaption></figcaption></figure>

<figure><img src="/files/hlwT5JK96VFKBOqLOkBT" alt="CLIP Text Encode"><figcaption></figcaption></figure>

## Workflow creation, saving, and downloading

<figure><img src="/files/RxplmlBeaegXb5I5Qrkg" alt="Workflow creation" width="314"><figcaption></figcaption></figure>

<figure><img src="/files/nl0MrYYbb7lUlu6ma7MF" alt="Publish workflow" width="391"><figcaption></figcaption></figure>


# 2-11 Canvas

Unleash your creativity with the SeaArt Canvas feature. Learn its basic functions and start making amazing AI art now.

> **SeaArt Canvas is an AI art tool that integrates multiple functions. It not only features basic painting editing functions such as brushes, layer management, and color adjustment but also incorporates advanced AI features including real-time generation, Text2Img, Img2Img,partial repaint, and AI Eraser. The comprehensive application of these features allows us to easily engage in diversified artistic creation. SeaArt Canvas offers a feature-rich yet user-friendly AI tool environment, enabling even first-time users of AI painting to quickly get the hang of it and complete the entire process from creative conception to image generation on this platform, greatly simplifying the creative process.**

<mark style="background-color:red;">More Advanced Tutorials</mark>

{% content-ref url="/pages/jW2ItNIkSvKCTVah0hB2" %}
[3-4 Canvas Guide](/guide-1/3-advanced-guide/3-4-canvas-guide)
{% endcontent-ref %}

## Familiar with Canvas

**Homepage entry:** Click on “Generate" at the top right corner - Canvas

<figure><img src="/files/mp1vAs6x29a612GyFBrI" alt="Canvas interface" width="563"><figcaption><p>Generate Canvas</p></figcaption></figure>

<figure><img src="/files/pAeVQtseAQcSAagk9nyI" alt="Introduction to functions of the Canvas interface" width="563"><figcaption><p>Realtime Canvas</p></figcaption></figure>

## Basic Function

### **real-time image creation**

Brush redraw

First, input your prompt words and utilize the Simple Pencil to draw corresponding elements on the canvas. You can adjust the redraw intensity to modify the final image effect, a higher intensity will result in more significant changes

<figure><img src="/files/XhwzAaZPDJu8v7n41U2N" alt="" width="341"><figcaption></figcaption></figure>

<figure><img src="/files/vL9qeSuykrEoNUWT8kyb" alt="Comparison  - brush drawing  and AI-generated cat image" width="375"><figcaption></figcaption></figure>

> Partial parameters:&#x20;
>
> Denoising Strength: 0.56&#x20;
>
> Preset Parameters: Graphic Design

Utilize preset parameters to achieve different image effects.

<figure><img src="/files/HQwInJl5o2TVlw3CmYzf" alt="Utilize preset parameters to achieve different image effects"><figcaption></figcaption></figure>

**real-time redraw**

Adding details and adjusting lighting effects.

You can also upload your own materials or use those provided by the platform. After positioning them appropriately, input your prompt words again and adjust the redraw intensity and preset parameters. This process adds more details and lighting effects to the image, enhancing its overall quality.

<figure><img src="/files/he5BfzzOdaucI95Orxr3" alt="Upload materials from computer or SeaArt" width="311"><figcaption></figcaption></figure>

<figure><img src="/files/i1nCdDkaHTzVBQLgGwjs" alt="Real-time redraw" width="563"><figcaption></figcaption></figure>

Additionally, you can use the Simple Pencil once more to add additional elements, which will be incorporated into the generated image based on the corresponding prompts.

<figure><img src="/files/EIEUaMmhscgPtUm4XJeZ" alt="Real-time redraw - add elements" width="563"><figcaption></figcaption></figure>


# 2-12 LoRA Training

Train your own LoRA models for AI art generation. Learn about dataset creation, image preprocessing, tagging, and publishing your trained LoRA. Let's dive in!

**Page entry:** Click on <mark style="background-color:yellow;">"Train"</mark> in the upper left corner of the homepage to access the style training page.

## **Operation Process**

1. Create Dataset: Upload dataset images, ensuring diversity in angles, lighting, backgrounds, etc., with a recommended count of around 30 images.
2. Image Preprocessing: Crop images, add taggers, and include trigger words.
   * Cropping Mod: Recommend focus cropping.
   * Size: Choose according to output requirements; Square: 512 \* 512; Portrait: 512 \* 768; Landscape: 768\*512.
   * Tagging Algorithm: Recommend Deepbooru.
   * Tagging Threshold: Recommend 0.6.
   * Trigger Words (Optional): A word that invokes Lora.
3. Choose preset parameters based on the trained Lora.
4. Enter model preview prompt words, mainly to review Lora's outputs during training.
5. Input the dataset name and click "Train Now".

<mark style="background-color:red;">\*For detailed instructions on Tagging and advanced training parameter settings, click here.</mark>

{% content-ref url="/pages/bDLwl9DD9y2ga18oYqz4" %}
[3-2  LoRA Training (Advance)](/guide-1/3-advanced-guide/3-2-lora-training-advance)
{% endcontent-ref %}

## Publish

1. Click on Train -> LoRA -> Publish to view the trained Lora.
2. Click on Publish, edit the Lora information, and fill in the relevant information as prompted, including LoraName, Tags (optional), Lora Cover, Model Usage, and Model Permissions.

Paid: When others use the model you published, you can receive corresponding Credits and get 70% of the Amount set.

<figure><img src="/files/lGJcxaJ1unNfsp7nnSHw" alt=""><figcaption></figcaption></figure>

3. Edit Version:

Includes: VersionName, Base Model, Trigger Words, Version Introduction.

<figure><img src="/files/qq8QEsvFiE8eZtkUpwKo" alt=""><figcaption></figcaption></figure>

4. Add image.

Upload the cover image for the LoRA and other showcase images.

5. Return to the LoRA page and click "Public."

## Edit the Lora again

Click on the Lora details page, then the three dots in the bottom right corner, and select "Edit.”

<figure><img src="/files/9hSkIPPnxsPndlvoxmZr" alt=""><figcaption></figcaption></figure>

## Take the Lora offline

After publishing the Lora, click on the Lora details page, then on the three dots in the bottom right corner, and select "Private.”

<figure><img src="/files/RfjZF63d5ftTs3OgbANn" alt=""><figcaption></figcaption></figure>

## &#x20;Training  the same Lora

1. Click on "Lora Training" on the homepage.
2. Select any Lora template, then click on "Train" in the top right corner.

<mark style="color:red;">\*</mark>You can modify parameters, datasets, etc., based on the same template.


# 3-Advanced Guide

Learn as we share tutorials for advanced SeaArt users. From Lora training to ComfyUI, explore the potential of SeaArt AI.


# 3-1 Principles of AI art

Learn the principles of AI art creation: from text prompts to image generation using text encoders, U-NET, and VAE.

> **The creative process: Enter prompts - Select model - Adjust relevant parameters - Generate image.**&#x20;
>
> **The process principle: Prompts - Text encoder - U-NET - VAE - Normal image.**

After entering the prompts, a text encoder processes them into embedding vectors. Simultaneously, a random seed generates a noise image. The parameters set for AI art, such as sampling method, sampling steps, image size, etc., are fed into the noise predictor. The noise predictor removes image noise, gradually revealing the image. Finally, the output is passed through a VAE to produce an image discernible to the naked eye.


# 3-2  LoRA Training (Advance)

Master AI art with advanced LoRA training! This guide covers everything from principles and processes to optimizing parameters for stunning, controllable results.

## The Principle of LoRA

#### What is LoRA used for?

Lora allows for fine-tuning the entire image while keeping the weights of the Checkpoint unchanged. In this case, only adjusting Lora is needed to generate specific images without modifying the entire Checkpoint. For some images that the AI has never encountered before, Lora is used for fine-tuning. This gives AI art a certain degree of "controllability.”

{% content-ref url="/pages/M9glU1B61lxg5D5YqggW" %}
[Image Training](/guide-1/3-advanced-guide/3-2-lora-training-advance/image-training)
{% endcontent-ref %}

Currently, the trained models are all "refinements" made on officially trained models (SD1.5, SDXL). Of course, refinements can also be made on models created by others.

{% content-ref url="/pages/b2qwnnM5Q3wKZM1gssgs" %}
[Video Training](/guide-1/3-advanced-guide/3-2-lora-training-advance/video-training)
{% endcontent-ref %}

Lora Training: AI first generates images based on the prompts, then compares these images with the dataset in the training set. By guiding AI to continuously fine-tune the embedding vectors based on the generated differences, the generated results gradually approach the dataset. Eventually, the fine-tuned model can produce results that are completely equivalent to the dataset, forming an association between the images generated by AI and the dataset, making them increasingly similar.

<mark style="color:red;">\*</mark>Compared to the Checkpoint, LoRA has a smaller file size, which saves time and resources. Moreover, it can adjust weights on top of the Checkpoint, achieving different effects.

####

## LoRA Training Process

> **Five steps: Prepare dataset - Image preprocessing - Set parameters - Monitor Lora training process - Training completion**

<mark style="background-color:yellow;">\*Taking the training of a facial Lora with SeaArt as an example.</mark>

#### Prepare the dataset

<mark style="color:red;">\*</mark>If you want to learn more about creating a dataset, you can read the guide below.

{% content-ref url="/pages/q3s7EznljaU20Cu4tK3i" %}
[How To Create Dataset For Training](/guide-1/3-advanced-guide/3-2-lora-training-advance/how-to-create-dataset-for-training)
{% endcontent-ref %}

When uploading the dataset, it's essential to maintain the principle of "diversified samples." This means the dataset should include images from different angles, poses, lighting conditions, etc., and ensure that the images are of high resolution. This step is primarily aimed at helping AI understand the images.

#### Image preprocessing

<mark style="background-color:yellow;">I. Cropping images II. Tagging III. Trigger words.</mark>

**I. Cropping images**

To enable the AI to better discern objects through images, it's generally best to maintain consistent image dimensions. You can choose from **512**\***512 (1:1), 512\*768 (2:3),** or **768\*512 (3:2)** based on the desired output.

Crop Mode: <mark style="background-color:yellow;">Center Crop / Focus Crop / No Crop</mark>&#x20;

Center Crop: Crops the central region of the image.

Focus Crop: Automatically identifies the main subject of the image.

<mark style="color:red;">\*</mark>Compared to center cropping, focus cropping is more likely to preserve the main subject of the dataset, so it is generally recommended to use <mark style="color:red;">focus crop.</mark>

**II. Tagging**

To provide textual descriptions for images in the dataset, allowing AI to learn from the text inside.

Tagging Algorithm: BLIP/Deepbooru

* BLIP: Natural language tagger, for example, "a girl with black hair."
* Deepbooru: Phrase language labels, for example, "a girl, black hair."
* Tagging Threshold: The smaller the value, the finer the description, recommended to be <mark style="color:red;">0.6.</mark>

Tagging process: Remove fixed features (such as physical features...) to allow AI to autonomously learn these features. Similarly, you can also add some features you want to adjust in the future (clothing, accessories, actions, background...).

<mark style="color:red;">\*</mark>For example, if you want all the generated images to have black hair and black eyes, you can delete these two tags.

**III. Trigger words**

Words that trigger the activation of Lora, effectively consolidating the character features into a single word.

<figure><img src="/files/wkWDXtSzWdya9Qvk7Qit" alt="Lora Training Process - Trigger Words"><figcaption></figcaption></figure>

#### Parameter Settings

**Base Model:** It is recommended to choose a high-quality, stable base model that closely matches the style of Lora, as this makes it easier for AI to match features and record differences.

**Recommended Base Models:**

Realistic: SD1.5, ChilloutMix, MajicMIX Realistic, Realistic Vison

Anime: AnyLoRA, Anything | 万象熔炉, ReV Animated

#### Advanced Config

<mark style="background-color:red;">**Training Parameters:**</mark>

<figure><img src="/files/otjYhjo8mY0870vi2o5U" alt="Lora Training Process - Training Parameters"><figcaption></figcaption></figure>

**Repeat (Single Image Repetitions):** The number of times a single image is learned. The more repetitions, the better the learning effect, but excessive repetitions may lead to image rigidity. <mark style="color:red;">Suggestion: Anime: 8; Realistic: 15.</mark>

**Epoch (Cycles):** One cycle equals the number of dataset multiplied by Repeat. It represents how many steps the model has been trained on the training set. For example, if there are 20 images in the training set and Repeat is set to 10, then the model will learn 20 \* 10 = 200 steps. If Epoch is set to 10, then the Lora training will have a total of 2000 steps. <mark style="color:red;">Suggestion: Anime: 20; Realistic: 10.</mark>

**Batch size:** It refers to the number of images the AI learns simultaneously. For example, when set to 2, the AI learns 2 images at a time, which shortens the overall training duration. However, learning multiple images simultaneously may lead to a relative decrease in the precision for each image.

Mixed precision: <mark style="color:red;">fp16</mark> is recommended.

<mark style="background-color:red;">**Sample Settings:**</mark>

<figure><img src="/files/OPe56mnITc3zkcqj5jao" alt="Lora Training Process - Sample Settings"><figcaption></figcaption></figure>

**Resolution:** Determines the size of the preview image for the final model effect.

* SD1.5: 512\*512

  SD1.5: 512\*512

**Seed:** Controls the randomly generated images. When using the same r seed with prompts, it will likely generate the same/similar images.

**Sampler \ Prompts \ Negative Prompts:** Mainly showcase the effect of the preview image of the final model.

<mark style="background-color:red;">**Save Settings:**</mark>

<figure><img src="/files/xvu7pl2w02Yd4gL3Z1wb" alt="Lora Training Process - Save Settings"><figcaption></figcaption></figure>

Determines the final number of Loras. If set to 2, and Epoch is 10, then 5 Loras will be saved in the end.

Save precision: Recommended <mark style="color:red;">fp16.</mark>

<mark style="background-color:red;">**Learning Rate & Optimizer:**</mark>

<figure><img src="/files/33n6Pl4knAgkWsASLHb1" alt="Lora Training Process - Learning Rate &#x26; Optimizer"><figcaption></figcaption></figure>

**Learning Rate:** It denotes the intensity of AI learning the dataset. The higher the learning rate, the more AI can learn, but it may also lead to dissimilar output images. When the dataset increases, it's advisable to try reducing the learning rate. It's recommended to start with the default value and then adjust it based on training results. It's suggested to gradually increase from a lower learning rate, recommended at <mark style="color:red;">0.0001.</mark>

**unet lr:** When the unet lr is set, the Learning Rate will not take effect. Recommended at <mark style="color:red;">0.0001.</mark>

**text encoder lr:** It determines the sensitivity to tags. Usually, the text encoder lr is set to <mark style="color:red;">1/2 or 1/10 of the unet lr.</mark>

**Lr scheduler:** It primarily governs the decay of the learning rate. Different schedulers have minimal impact on the final results. Generally, the default "cosine" scheduler is used, but an upgraded version, <mark style="color:red;">"Cosine with Reastart,"</mark> is also available. It goes through multiple restarts and decays to fully learn the dataset, avoiding interference from "local optimal solutions" during training. If using "Cosine with Reastart," set the Restart Times to <mark style="color:red;">3-5.</mark>

**Optimizer:** It determines how AI grasps the learning process during training, directly impacting the learning results. It's recommended to use <mark style="color:red;">AdamW8bit.</mark>

<mark style="color:red;">Lion:</mark> A newly introduced optimizer, typically with a learning rate about 10 times smaller than AdamW.

<mark style="color:red;">Prodigy:</mark> If all learning rates are set to 1, Prodigy will automatically adjust the learning rate to achieve the best results, suitable for <mark style="background-color:yellow;">beginners.</mark>

<mark style="background-color:red;">**Network:**</mark>

<figure><img src="/files/PvihDPG66TjOPZ0TnfdA" alt="Lora Training Process - Network"><figcaption></figcaption></figure>

Used to build a suitable Lora model base for AI input data.

**Network Rank Dim:** It directly affects the size of Lora. The larger the Rank, the more data needs to be fine-tuned during training. 128=140MB+; 64=70MB+; 32=40MB+.

**Recommended:**

<mark style="color:red;">Realistic: 64/128</mark>

<mark style="color:red;">Anime: 8/16/32</mark>

Setting the value too high will make the AI learn too deeply, capturing many irrelevant details, similar to "overfitting”

**Network Alpha:** It can be understood as the degree of influence of Lora on the original model weights. The closer it is to Rank, the smaller the influence on the original model weights, while the closer it is to 0, the more pronounced the influence on the original model weights. Alpha generally does not exceed Rank. Currently, Alpha is typically set to <mark style="color:red;">half of Rank</mark>. If set to 1, it maximizes the influence on weights.

<mark style="background-color:red;">**Tagging Settings:**</mark>

<figure><img src="/files/Z1Gj3pM785h4zdLqJVvK" alt="Lora Training Process - Tagging Settings"><figcaption></figcaption></figure>

In general, the closer a tag is to the front, the greater its weight. Therefore, it's usually recommended to enable <mark style="color:red;">Shuffle Caption</mark>

## LoRA Training Issues

### **Overfitting / Underfitting**

**Overfitting:** When there is a limited dataset or the AI matches the dataset too precisely, it leads to Lora generating images that largely resemble the dataset, resulting in poor generalization ability of the model.

<figure><img src="/files/asBQW499RCpxGwT2jb12" alt="Lora Training Process - Training issues - Overfitting"><figcaption></figcaption></figure>

The image on the top right closely resembles the dataset on the left, both in appearance and posture.

**Reasons for Overfitting:**

1. The dataset is lacking.
2. Incorrect parameter settings (tags, learning rate, steps, optimizer, etc.).

**Preventing Overfitting:**

1. Decrease learning rate appropriately.
2. Shorten the Epoch.
3. Reduce Rank and increase Alpha.
4. Decrease Repeat.
5. Utilize regularization training.
6. Increase dataset.

**Underfitting:** The model fails to adequately learn the features of the dataset during training, resulting in generated images that do not match the dataset well.

You can see that Lora's generated images fail to adequately preserve the features of the dataset — they are dissimilar.

**Reasons for Underfitting:**

1. Low model complexity
2. Insufficient feature quantity

**Preventing Underfitting:**

1. Increase learning rate appropriately
2. Increase Epoch
3. Raise Rank, reduce Alpha
4. Increase Repeat
5. Reduce regularization constraints
6. Add more features to the dataset (high quality)

### Regular Dataset

A way to avoid overfitting of images is by adding additional images to enhance the model's generalization ability. The regular dataset should not be too extensive, otherwise, the AI will overly learn from the regular dataset, leading to inconsistency with the original target. It is recommended to have <mark style="color:red;">10-20</mark> images.

For example, in a portrait dataset where most images feature long hair, you can add images with short hair to the regular dataset. Similarly, if the dataset consists entirely of images with the same artistic style, you can add images with different styles to the regulardataset to diversify the model. The regular dataset does not need to be tagged.

<mark style="color:red;">\*</mark>In layman's terms, training Lora in this way is somewhat like a combination of the dataset and a regular dataset.

### Loss

The deviation between what AI learns and reality, guided by loss, can optimize the direction of AI learning. Therefore, when the loss is low, the deviation between what AI learns and reality is relatively small, and at this point, AI learns the most accurately. As long as the loss gradually decreases, there are usually no major issues.

The loss value for Realistic images generally ranges from <mark style="color:red;">0.1 to 0.12</mark>, while for anime, it can be lowered appropriately.

Use the loss value to assess model training issues.

<figure><img src="/files/SUs6Nszn0a5mmq5SzEQz" alt="Lora Training Process - Loss Value to Assess Model Training Issues"><figcaption></figcaption></figure>

### Summary

Currently, the "fine-tuning models" can be roughly divided into three types: the Checkpoint output by Dreambooth, the Lora, and the Embeddings output by Textual Inversion. Considering factors such as model size, training duration, and training dataset requirements, Lora offers the best "cost-effectiveness". Whether it's adjusting the art style, characters, or various poses, Lora can perform effectively.

<figure><img src="/files/ruoRYLhj8kxp7M92P8Pk" alt="Checkpoint ReV Animated &#x26; Invisible Concept Lora"><figcaption></figcaption></figure>

## SDXL LoRA Setting

### Recommended Training Parameters:

### **Epochs and Repeats**

#### Epochs:

Number of dataset image training cycles. <mark style="background-color:yellow;">We suggest 10 for beginners.</mark> The value can be raised if the training seems insufficient due to a small dataset or lowered if the dataset is huge.

#### Repeats:

Number of times an image is learned. Higher values lead to better effects and more complex image compositions. Setting it too high may increase the risk of overfitting. Therefore, <mark style="background-color:yellow;">we suggest using 10</mark> to achieve good training results while minimizing the chance of overfitting.&#x20;

<mark style="background-color:yellow;">Note: You can increase the epochs and repeats if the training results do not resemble.</mark>

### Learning Rate and Optimizer:

#### learning\_rate (Overall Learning Rate):

Degree of change in each repeat. Higher values mean faster learning but may cause model crashes or inability to converge. Lower values mean slower learning but may achieve optimal state. This value becomes ineffective after setting separate learning rates for U-Net and Text Encoder.

#### unet\_lr (U-Net Learning Rate):

U-Net guides noise images generated by random seeds to determine denoising direction, find areas needing change, and provide the required data. Higher values mean faster fitting but risk missing details, while lower values cause underfitting and no resemblance among generated images and materials. The value is set accordingly based on the model type and dataset. We suggest 0.0002 for character training.

#### text\_encoder\_lr (Text Encoder Learning Rate):

It converts tags to embedding form for U-Net to understand. Since the text encoder of SDXL is already well-trained, there is usually no need for further training, and default values are fine unless there are special needs.

#### Optimizer:

An algorithm in deep learning that adjusts model parameters to minimize the loss function. During neural network training, the optimizer updates the model's weight based on the gradient information of the loss function so the model can better fit the training data. The default optimizer, AdamW, can be used for SDXL training, and other optimizers, like the easy-to-use Prodigy with adaptive learning rates, can also be chosen based on specific requirements.

#### lr\_scheduler (Learning Rate Scheduler Settings):

Refers to a strategy or algorithm for dynamically adjusting the learning rate during training. Choosing Constant is sufficient under normal circumstances.

### Network Settings:

#### network\_dim (Network Dimension):

Closely related to the size of the trained LoRA.&#x20;

For SDXL, a 32dim LoRA is 200M, a 16dim LoRA is 100M, and an 8dim LoRA is 50M. For characters, selecting 8dim is sufficient.

#### network\_alpha:

Typically set as half or a quarter of the dim value. If the dim is set as 8, then the alpha can be set as 4.

### **Other settings:**

#### Resolution:

Training resolution can be non-square but must be multiples of 64. For SDXL, we suggest 1024*1024 or 1024*768.

#### enable\_bucket (Bucket):

If the images' resolution is not unified, please turn on this parameter. It will automatically classify the resolution of the training set and create a bucket to store images for each resolution or similar resolution before the training starts. This saves time on unifying the resolution in the early stage. <mark style="background-color:yellow;">If the images' resolution has already been unified, there is no need to turn it on.</mark>

#### **noise\_offset and multires\_noise\_iterations:**

Both noise offsets improve the situation where the generated image is overly bright or too dark. If there are no excessively bright or dark images in the training set, they can be turned off. If turned on, <mark style="background-color:yellow;">we suggest using multires\_noise\_iterations</mark> with a value of 6-10.

#### multires\_noise\_discount:

Needs turning on with the multires\_noise\_iterations mentioned above, and a value of 0.3-0.8 is recommended.

#### clip\_skip:

Specifies which text encoder layer's output to use counting from the last. Usually, the default value is fine.


# How To Create Dataset For Training

A quality dataset is key to effective model training. This guide covers how to find or create one for different model types.

## What Is The Dataset?

Training Dataset is source of our models, it's essential to use correct size, tagging and edit to create good dataset.

* Usually 20\~40 images are enough for many type of LoRAs but Checkpoint (Model) training require a lot more image (minimum 100, which is max of SeaArt's dataset), so it's recommended to only focus on training LoRAs at the moment to save our time and credits. If you are going to create detailed LoRA you can prefer to use more images but remind that, **More Images Doesn't Mean Better LoRA.** It's best option to choose various images with same content for our LoRAs.
* Dataset can be found in web images but also can be created via AI image generation. Dataset must have no wrong, bad part so it's kinda tricky to use AI to create dataset. However, if you trust your skills and luck you are free to create dataset with image generation.

## Selecting/Creating Images For Dataset

There is a different type LoRAs and they require different type of images but there is a few important point which applies to all datasets.

### Can i use something can't be generated with AI?

Images should be something can be generated on AI. (Doesn't include main subject of LoRA) For example, you want to create a character LoRA of a game character. It's okay to can't generate that character with AI but you must be able to generate that spesific image with AI for best possible results. Character: Frieren (can't be generated with model you want to use.) Dataset must be images AI could generate if it knew who is Frieren.&#x20;

**Examples**

&#x20;`Image of Frieren standing and smiling to viewer`

* Ai can create standing and smiling characters even if it's not Frieren, so image can be used.

`Portrait of Frieren with sad expression`

* Ai can create portrait of a woman with sad expression even it's not Frieren, so image can be used.

`Image of Frieren's upper body stucked in mimic(chest monster)`

* Many of models can't create proper image of an human got stuck in a mimic even if it's not Frieren, so image can't be used for models that can't generate it.

Let's say i want to create a style LoRA, then style is not problem if AI can't generate it but it must be able to generate subject.&#x20;

Style: Dark, gothic style (can't be generated with model you want to use.)

**Examples**

`Illustration of a castle with dark, gothic style`

* Illustration of a castle can be generated with AI even if it's not style you want, so image can be used.

`Image of a half apple, half cat armored cyborg with dark, gothic style`

* Many models fail at generating half apple, half cat armored cyborg even it's not style we want, so image can't be used.

It can sound absurd but it's very important for correct tagging and best LoRAs comes with correctly tagged datasets. Main point is, Image must doesn't include any confusing, hard to generate part in itself unless it's not the subject of your LoRA.&#x20;

### Selection of Dataset for Character LoRA

* Your character's style must be similar to model you want to use. For example:

Anime style Naruto LoRA on Anime Model: Dataset must be anime style images. Photorealistic Luffy LoRA on Photorealistic Model: Dataset must include realistic images of subject not anime style.

* Usually **25-40** image of character is enough, if you are not going to create a LoRA with a lot of detail. Using a lot of image can cause overcooked LoRAs.
* Use same character images with a few different poses, angles, views, clothes, expressions, background etc. and combine them to create various images.

**Example**&#x20;

**Character:** Luffy from One Piece&#x20;

**Pose:** Standing, sitting, action poses (could be waving hand, peace sign with hand, wide smile with closed eyes or anything you desire but must be something can generated on AI)&#x20;

**Angle/View:** Front view, upper-body shot, full-body shot, portrait, back view etc.&#x20;

**Clothes:** Original clothes of Luffy, formal wear, t-shirt and jeans, suit etc.&#x20;

**Expressions** (recommended to use in only portraits or close views): Smile, angry, sad, confused, shy etc.&#x20;

**Background:** City, beach, sea, ship, home etc.

**Images**

* upper-body shot, Luffy, smiling, standing, formal wear, beach
* upper-body shot, Luffy, sad, crying, waving to viewer, original clothes, ship
* full-body shot, Luffy, sitting, t-shirt and jeans, home ...

After finding/creating spesific images with combination of different terms, now you have a good dataset with various images. It increase flexibility and generation capability of LoRA.&#x20;

If you use portrait in all images, it will push portrait shots and cause deformations in different angles. If you use original clothes of Luffy in every image, it will tend to create straw hat, red jacket like clothes everytime even you don't want to.&#x20;

If you use beach as a background in every image, it will generate images with beach background even if you don't ask for it.

* Sweetpoint is using a term **max** in 3-6 images. Angle/View can be used more, for example:

Dataset with 30 image: 12 portrait shot, 4 beach background, 4 smiling expression, 6 image with original clothes, 4 backview, 6 upper body, 8 full body, 6 sitting, 8\~10 standing etc.

* Try to find/create different combinations, don't use similar tags for one term. If you create six image of `Luffy with orijinal clothes` and all of them are `back view` then, either you can get `back view` while trying to create him with original clothes or you can get his clothes while generating `back view` such as red jacket, straw hat even if you don't ask for it.

### Selection of Dataset for Style LoRA

* Base model you are going go use must be flexible or <mark style="background-color:yellow;">similar</mark> to style you want to create.
* Subject must be something AI can create.
* Use various images with different subjects but <mark style="background-color:yellow;">same style.</mark>
* <mark style="background-color:yellow;">30-40</mark> image is enough for a style LoRA.
* <mark style="background-color:yellow;">15-20 image,</mark> if style of your images are confident and similar. Can be used for pixel art like styles.
* <mark style="background-color:yellow;">+50 image,</mark> if you are going to make very detailed flexible style LoRA that can be used on any subject. (Pick different subjects as much as possible)

**Example for 30-40 image dataset, values can be adjusted depending on your total image count.**&#x20;

<mark style="color:red;">Note:</mark> All images must have <mark style="color:red;">same style</mark> (desired style to create LoRA)&#x20;

Image counts are suggested by myself, depending on what people prioritize and generate with AI. You can change values for different purposes but if you are going to make standart style LoRA, i highly recommend to use similar image count.

* 10-14 woman image (different characteristics, angles, views, poses, clothes, background)
* 8-10 man image (different characteristics, angles, views, poses, clothes, background)
* 6 landscape/scenery image with atleast 3 different place (without any main object)
* Various objects with different views, can be an fruit, car, house, pen, fire or anything you desire. Important part is having same style in images. **Good for big datasets to increase flexibility of LoRA.**

### Crop Mode

Crop is important to make our dataset in correct resolution. There is 3 different option for crop mode and they have different effects.&#x20;

**Center Crop:** Take center of image as a reference point and crop your images to correct size.&#x20;

**Focus Crop:** Take main subject as a reference point and crop your images to correct size.&#x20;

**No Crop:** Doesn't crop your images, it can be used if images are already in correct size.

* <mark style="background-color:yellow;">Focus Crop</mark> is recommended for character LoRAs, if images are not in correct size.

### Resolution

* `512x512, 512x768, 768x512` is recommen ded for Stable Diffusion 1.5 LoRAs. Although, you can use `768x768` if you want to create highly detailed LoRA.
* `1024x1024` is recommended for SDXL, Flux and Stable Diffusion 3.5 training.

## Tagging The Dataset

Tagging is very important for better LoRAs, it can take some of your time but results will be worth of your time.&#x20;

**How tagging works and what does it mean?**&#x20;

It can look confusing in first sight but it's actually pretty similar to how we prompt in image generation. To explain it basically;

* We have a image in dataset but AI doesn't know what is it.
* We add tags to image in dataset, basically we create a prompt without any order so <mark style="background-color:yellow;">AI understands and learns what prompt (tags) created that image.</mark> Also that's why things can be generated with AI is recommended because if AI is not capable of generating that image with given tags, you will just teach AI to something it can't generate and it cause a lot of wrong/deformed generation.
* After that, AI learn that tags and image to be able create similar results with its dataset.

### Tagging Algorithm

We can set tags by ourselves or use a tag algorithm to make it faster and easier for us. There is two different tagging algorithm we use in creating dataset.

#### BLIP

This system is more similar to natural language, it uses small and incomplete sentences to identify our images. It's recommended to use it on SDXL based models and Photorealistic style models. It's best option for <mark style="background-color:yellow;">Flux and SD3.5</mark> trainings but tagging them by yourself is probably better, if you have enough time and patience.

#### Deepbooru

A tagging system that uses `booru` website (Danbooru, Safebooru, e621 etc.) tags to identify our images. They are certain terms that used for tagging these website sharings and have pretty big database. It's recommended to use it on <mark style="background-color:yellow;">SD1.5</mark> based LoRA creations and especially on <mark style="background-color:yellow;">Anime/Furry/Cartoon based models.</mark>

#### Tagging Threshold

It can be used between 0-1, adjusts the tag count and description level of algorithm.

* Lower values create more tag and aims to describe every detail of image, it sounds good but it add tons of unnecessary tag to database and it's something we don't want.
* Higher values create less tag and only identify general details of image, it can cause bad results to have undefined details in tagging process.
* Using it between <mark style="background-color:yellow;">`0.5-0.8`</mark> is recommended depending on what you make and how much detailed you want to tag the dataset images.

## Editing Tags And Usage of Trigger Word

Unfortunately auto-tagging system is not 100% correct so we have to edit our database with removing unnecessary tags and add important tags. Most essential part of LoRA is `trigger words` because they include the data we want to teach to AI.

### Trigger Word

**Trigger Word must be used for best working LoRAs, it's very important part of prompting phase.**&#x20;

Trigger word is activation key of LoRAs, it's basically tags doesn't written to data of image but shown in image. Consistent parts of LoRA must be identified by trigger words not with their own tag. For example, our character has a yellow hair and we want to make LoRA of her. Tag data must not include `yellow hair, blonde` like terms so AI can take them as a part of trigger word.

### How Should I Edit Tags?

* First thing to do is removing unnecessary tags from our LoRA, it can be tag of an small detail like <mark style="background-color:yellow;">mole, necklace</mark> etc. It's mostly depend on what you prioritize and try to create.
* Secondly, we have to remove tags of consistent details. If we want to have blonde character LoRA, then remove <mark style="background-color:yellow;">yellow hair, blonde</mark> like terms. If we want to get an ink illustration style LoRA, then remove tags related to <mark style="background-color:yellow;">ink, illustration, art style</mark> etc.
* Last step, add important tags that describe image but they must be something not related to your LoRA's main subject.

For example, add tags related to <mark style="background-color:yellow;">background, lighting, blur, pose, expression</mark> that describes image for **a character LoRA,** or add tags that describe main subject of image with other details for **a style LoRA.**

## End of Guide

Now your dataset has correctly prepared images with important tags that describe image as a prompt does so AI will understand what created that images.&#x20;

Also your `trigger word` will be placeholder of your LoRAs main subject (missing tags of images you used), with this way AI will learn what does your `trigger word` means and educate itself on this topic.

**Example**

&#x20;LoRA: Naruto Character LoRA

&#x20;Image: Naruto standing in front of a tree and smiling&#x20;

Tags: 1boy, standing, tree, smiling etc.&#x20;

Trigger Word: Naruto&#x20;

Trigger word is placeholder of all spesific characteristics of Naruto such as <mark style="background-color:yellow;">yellow hair, blue eyes, cheek marks</mark> etc.&#x20;

Don't add these type characteristics to tags so AI will consider your `Trigger Word` as a activation key of these features.&#x20;

That's why we actually train LoRAs, to make our `Trigger Word` an activation key and meaningful word.


# Image Training

## Concept of LoRA

* **LoRA** allows fine-tuning specific image features **without modifying the base Checkpoint weights**. This means you can generate targeted image results just by adjusting the LoRA, instead of retraining the entire model.
* Currently, LoRA training is typically conducted on official base models such as **SD1.5, SDXL, PONY, ILLUSTRIOUS, FLUX, SD3.5**. Refinement can also be performed on community-created models.
* **How LoRA training works:** The AI first generates images based on prompts. These are then compared with the images in your dataset. The system gradually adjusts the **embedding vectors** based on the differences, guiding the AI to produce results increasingly similar to the dataset. Eventually, the model can generate images nearly identical in style or subject to the dataset, building a strong associative connection.

### LoRA Training Workflow

**Five steps:**\
**Prepare dataset → Image preprocessing → Set parameters → Monitor training → Complete training**

#### Dataset

High-quality datasets are key to effective model training. Training datasets are the source of our models, and using the correct size, tagging, and editing is crucial for creating good datasets.

● Usually, 20 to 40 images are sufficient for many types of LoRA, but style model training requires more images for better generalization. More images aren't always better; adding low-quality materials can reduce model quality.

● Materials can typically be found on image websites or created using AI image generation. But low-quality materials should be avoided in both cases.

Low-quality materials generally have these characteristics: scenes too dark to identify, unreasonable composition, unclear images, details difficult to replicate, unclear primary/secondary elements, details difficult to stack repeatedly, appearance of irrelevant elements, inconsistent subjects with image splicing, incorrect or misaligned limbs.

● The character style being trained should ideally be similar to the chosen Checkpoint style, i.e., anime characters should use anime Checkpoints for training.

Usually, 25-40 images of the same character are sufficient. Use the same character images but with different poses, angles, views, clothes, expressions, backgrounds, etc.

A dataset containing 30 images: 12 portrait photos, 4 beach background photos, 4 smiling expression photos, 6 photos wearing original clothes, 4 back view photos, 6 upper body photos, 8 full body photos, 6 sitting pose photos, 8-10 standing pose photos, etc.

● LoRA Style Dataset Selection

● The benchmark model used must be flexible or similar to the style being trained; illustration styles should use flat-type Checkpoints for training.

All images must have the same style (the style needed to create the LoRA).

#### Editing Tags and Trigger Words

● Trigger words must be used for LoRA to work optimally; they are a very important part of the prompting phase.

● Trigger words are the activation keys for LoRA, essentially tags that aren't written into the image data but are displayed in the image. The consistent parts of LoRA must be identified through trigger words, not through its tags.

For example, if character training has fixed hair features, there's no need to note this in the tags, allowing these features to merge into the trigger word. For style training, style vocabulary tags need to be deleted.

#### Tagging

| Name        | wd1.4                | deepbooru           | blip                      | joy2                      | llava                     |
| ----------- | -------------------- | ------------------- | ------------------------- | ------------------------- | ------------------------- |
| Output Form | Words                | Words               | Natural Language          | Natural Language          | Natural Language          |
| Main Use    | Anime Specialization | Anime Image Sorting | General Image Description | General Image Description | General Image Description |
| Model Usage | 1.5, il, pony        | 1.5, il, pony       | flux, xl                  | flux, xl                  | flux, xl                  |
| Threshold   | 0.3-0.6              | 0.3-0.6             | /                         | /                         | /                         |

Threshold: Lower values mean more detailed descriptions.

#### Image Preprocessing

Cropping Images

● Center cropping: Crops the center area of the image.

●Focus cropping: Automatically identifies the main subject of the image.

● No cropping: No image cropping, must be used with ARB bucketing.

● Compared to center cropping, focus cropping more easily preserves the subject of the dataset, so focus cropping is generally recommended.

● For Stable Diffusion 1.5 LoRA, 512x512, 512x768, and 768x512 are recommended. If you want to create highly detailed LoRA, 768x768 can also be used.

● For SDXL, Flux, and Stable Diffusion 3.5 training, 1024x1024 is recommended.

#### Dataset Creation and Upload

You can choose to upload existing datasets (i.e., the collective term for corresponding images and text annotations) or upload images for tagging and cropping processing.

Uploaded datasets cannot be recropped or automatically tagged, but tags can be manually modified.

After uploading images (batch upload supports up to 50 images at a time), you can select the cropping method, size, and tagging method. After selection, click Crop/Tag (wait until processing is complete to start adjusting parameters for training).

#### Training Parameter Settings

● At the top is the base model type (i.e., the large model type under which we train LoRA).

● Base Model: Different choices for each type of base model.

● Repeat: How many times each image is trained.

● Epoch: How many cycles all images are trained.

● Model Effect Preview Prompts: After training is complete, each model will have a sample image; the prompt used to generate this sample image is the model effect preview prompt.

<figure><img src="/files/ymYkknxjXOYcHxjODj9k" alt=""><figcaption></figcaption></figure>

#### Advanced Parameter Settings

● Batch size: Refers to the number of data samples sent to the model at once. When set to 4, it means the model processes 4 images each time. By processing data in batches, memory utilization and training speed can be improved. Typically, Batch Size values are chosen as powers of 2. Increasing Batch Size allows for a proportional increase in learning rate, e.g., 2 times the batch\_size can use twice the UNet learning rate, but the TE learning rate cannot be increased too much.

● Gradient Checkpointing: A training algorithm that trades computation for VRAM, saving memory but sacrificing some speed. If Batch Size is 1, it's turned off; if Batch Size is 2 or above, it's turned on.

● ARB Bucketing: Used to train with images of non-fixed aspect ratios (with ARB bucketing enabled, no cropping is needed. ARB bucketing will increase training time to some extent. ARB bucket resolution must be greater than training material resolution).

● ARB Bucket Minimum Resolution: Default is 256; uploaded image resolution cannot be less than 256.

● ARB Bucket Maximum Resolution: Default is 1024; uploaded image resolution cannot be greater than 1024. You can increase the value to add materials with a greater resolution.

● ARB Bucket Resolution Steps: Default is 64, which is fine.

● Save Every N Epochs: Saves models based on cycle count. E.g., determines the final number of LoRAs saved. If set to 2, and Epoch is 10, then 5 LoRAs will be saved in the end.

● Learning Rate: It represents the intensity with which AI learns the dataset. Higher learning rates mean stronger AI learning ability, but may also lead to inconsistent output images. It's recommended to gradually increase from lower learning rates; the suggested learning rate is 0.0001.

● unet lr: When unet lr is set, the learning rate will not take effect. The recommended setting is 0.0001.

● text encoder lr: Determines sensitivity to tags. Typically, it is set to 1/2 or 1/10 of the unet lr.

● Learning Rate & Optimizer:

|               | AdamW8bit                                                                         | prodigy                                                                              |
| ------------- | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ |
| Learning Rate | <p>Total learning rate 1e-4</p><p>Scale up proportionally based on batch size</p> | <p>All learning rates set to 1</p><p>Actual learning rate will adjust adaptively</p> |
| Lr scheduler  | <p>Cosine with restart</p><p>Restart count not exceeding 4</p>                    | constant                                                                             |
| Lr warm up    | Warm-up steps are 5%-10% of total steps                                           | /                                                                                    |

● Network: Common values

| Network Rank Dim | 32 | 64 | 128 |
| ---------------- | -- | -- | --- |
| Network Alpha    | 16 | 32 | 64  |

Setting this value too high will cause AI to learn too deeply and make the model larger, capturing many irrelevant details, similar to "overfitting."

● Shuffle Caption: When enabled, the token order of the text will be randomly shuffled during training to enhance the generalization ability of the generation model. It is recommended to turn it on.

● Keep N Tokens: Generally choose 1, to keep our first entered trigger word with the highest weight.

● Noise Offset: Adds global noise during training, improving the brightness range of images (meaning it can generate darker or whiter images).

● Multires Noise Iterations: Defines the number of iterations.

● Multires Noise Discount: Defines the proportion by which noise gradually decreases with iterations.

Note: As they (Noise Offset, Multires Noise Iterations, Multires Noise Discount) all require extra steps to ensure convergence, the training time will be affected.

#### View Training Records

View training records on the right side of the dataset creation screen.

<figure><img src="/files/DsWMV1RQNIJtdvaMwn7K" alt=""><figcaption></figcaption></figure>

Select models with no issues in the sample images and save them.

Click on your account avatar in the upper right corner to jump to the personal work screen, and select the Model tab to use.

#### Model Testing

SeaArt will automatically synchronize the base model and the trained LoRA

Turn on the Fixed Seed number in the Advanced Config, and adjust the model weight to test the effect of the model at different weights.

#### LoRA Training Issues

Overfitting/Underfitting

Overfitting: When the dataset is limited or the AI matches the dataset too precisely, the LoRA generates images very similar to the dataset, resulting in poor generalization ability.

<figure><img src="/files/lVRlWX1rFvsDcr4mGvJS" alt=""><figcaption></figcaption></figure>

The image in the upper right is very similar to the dataset on the left in appearance and pose.

Causes of Overfitting:

● Lack of dataset

● Incorrect parameter settings (tags, learning rate, steps, optimizer, etc.)

Preventing Overfitting:

● Appropriately reduce learning rate.

● Reduce Epoch.

● Reduce Repeat.

● Use regularization training.

● Increase dataset.

Underfitting: The model fails to adequately learn the features of the dataset during training, resulting in generated images that don't match the dataset well.

Causes of Underfitting:

● Low model complexity

● Insufficient features

Preventing Underfitting:

● Appropriately increase learning rate.

● Increase Epoch.

● Increase Repeat.

● Reduce regularization constraints.

● Add more feature materials (high quality) to the dataset.

#### Regular Dataset

One way to avoid image overfitting is to add additional images to enhance the model's generalization ability. Regular datasets should not be too large, otherwise the AI will over-learn the regular dataset, leading to inconsistency with the original target. 10-20 images are recommended.

For example, in a portrait dataset where most images feature long hair, add short hair images to the regular dataset. Similarly, if the dataset consists entirely of images in the same artistic style, add images of different styles to the regular dataset to enrich the model. Regular datasets don't need to be tagged or cropped.

#### Image Model Type Classification

● SD1.5: Released in October 2022. The mainstream training size is 512\*512. As a training base model, the training speed is fast, but the image quality is relatively average.

● Common models include:

● SDXL: Released in July 2023. The mainstream training size is 1024\*1024. As a training base model, the training speed is average, and the image effect is better.

● Common models include:

● Pony: There are many versions, and the V6 XL version released in January 2024 is the most popular. The mainstream training size is 1024\*1024. Focuses on the cartoon and animal-style image generation.

● Common models include:

● Illustrious: The V1.0 version released in July 2024 is the most popular. The mainstream training size is 1024\*1024. Focused on providing high-quality anime and illustration style image generation capabilities.

● Common models include:

● Note: Models trained on Illustrious and Pony base models can be used under the SDXL models.

● Flux: Released in August 2024, based on a novel transformer architecture, using 12 billion parameters, allowing it to generate detailed and realistic images.

The mainstream training size is 1024\*1024, but 512\*512 also produces good effects.

● Model versions include: FLUX Pro, FLUX Dev, FLUX Schnell, FLUX GGUF, NF4

● Common models include:

### Wan 2.2 Text-to-Image Training

The table shows the impact of low-noise and high-noise LoRA on the final generated images.

From the table we can see that whether you are training a character or a style, you must train a low-noise model.

The factors affected by the low-noise model are very important. The low-noise model affects the character's facial features and the expression of the style.

The high-noise model affects composition and lighting/shadows.

So there are only two training combinations: Low Noise, Low Noise + High Noise

* ### Realistic Portrait Training

Use the Flux model as the base model.

Add trigger words and choose deepseek to perform annotation.

Next, check whether the caption contains errors.

For example: She appears to be looking up at the waterfall with an expression of awe. Anything that doesn't match the image content should be corrected. Vague word descriptions, such as "appears" or "possibly," should also be corrected.

Correct it to: She looks up at the sky, her mouth agape, her expression one of profound reverence.

Next, you need to clearly indicate which part of the image the trigger-word character refers to (if you do not modify this, for Flux, it may not affect that much, but if you switch to Qwen-image, the character may end up looking completely different from the person in the dataset images).

For example: The woman is wearing a white dress and standing near a waterfall.

Correct it to: The woman named tts is wearing a white dress and is standing near the waterfall.

After all checks and corrections are completed, adjust the parameters.

Then fill the corrected caption into Model Effect Preview Prompts.

Whether it is on or off, the Enable ARB Bucketing option has no impact, as all images in the dataset are 1024\*1024, which is within the ARB Bucket Maximum Resolution.

![descript](/files/DFMOVVxcpj4IiFiybcIC)

Set Learning Rate Scheduler to constant.

Set Optimizer to Prodigy.

When using Prodigy, there is no need to manually adjust the learning rate, as this optimizer uses a fully adaptive learning rate.

![descript](/files/NHuV0IQu0LLPG7TPh4px)

For Network Size and Network Alpha, you do not need to set them too high for portrait training.

When training portraits with Flux using natural language, there is no need to enable Shuffle Caption.

![descript](/files/A4HSXiNdpWMZBLimCOSa)

If the number of steps is relatively small, you can skip adding noise.

![descript](/files/imfOV2E3xZLrPrj63yAR)

![descript](/files/oRdjAkU15k6RWo3mtbjJ)

Keep the results from the last few epochs.

Use many different scenes to test image generation.


# Quick Training Guide

#### Training Entry

From the Create dropdown menu, select Model -> Training to enter the dataset screen.

Click Train Now to enter the training screen.

You can also click the training icon in the lower left corner of the creation flow to enter.

<figure><img src="/files/f04PraBMrWnhCmwXE7Tg" alt=""><figcaption></figcaption></figure>

Basic Training Process

#### Image Material Selection

Upload image resolution: 256\*256 ≤ image size ≤ 2048\*2048

Supported image formats: .png .jpg .webp

Note: The quality of image materials determines the final effect of the model.

Number of training images: 10-30 images are recommended; the more images uploaded, the better the effect.

#### Selection Key Points

The main points for selecting images are displayed at the bottom of the screen.

**Clear Character**

This means the character's facial features are clear without obstruction, the scene does not cause visual errors for the character, and the character occupies the main part of the image.

**Rich Angles**

**This refers to different views of the character, such as front, side, back, from below, and from above. It also includes various poses, such as standing, squatting, and lying down, as well as different expressions.**

**Rich Backdrop**

This refers to images of the character in various scenes, including different decorations (clothing, hair accessories, expressions), different lighting, as well as full-body shots, half-body shots, actions in different scenes, close-ups of parts, such as facial features, hands, feet, etc. These images should account for about 30% of the total.

**Varied Styles**

When training realistic characters, you cannot select anime-style images or images in other styles, and vice versa.

**Blurry Image**

Image materials should be high-definition and of good quality.

**Watermark**

Images cannot have watermarks. Otherwise, the final model will likely also generate images with watermarks, affecting the model's effect.

#### Training Steps

Click the pencil to modify the training set's name.

Training Type: Currently, only character is available (other types will be enabled later).

Training Theme: Divided into Realistic Character and Anime Character (including furry).

Training Model: You can select various types of models.

Click Upload Images on the right to upload images from your local device. The box on the left is for uploading personal works (you can generate character images through the creation flow for uploading).

After uploading some images, you can also select Upload Image from above to upload more.

If there are upload errors, you can click Delete All on the left to delete images.

The Training Strength in the lower left is determined based on the number of images, and you can select the strength according to the recommendation.

The icon in the upper right is used to save the uploaded images.

Use the Training Object Name field to give your trained character a name, making it more convenient for you to generate this character after completing training.

Inconsistent upload image resolutions will not affect the model effect.

Click Train Now to start training.

#### View Training Test

You can quickly switch to the task list by clicking the icon next to the Train Now button in Quick Training.

In the dataset screen, click Task List to enter the training queue, then click View Training to check the training status.

<figure><img src="/files/UOKKGpQsNy76x7bxU54B" alt=""><figcaption></figcaption></figure>

Here you can check the training progress or stop the training.

<figure><img src="/files/Q1iQA8QmESaYKv2m3P14" alt=""><figcaption></figcaption></figure>

Select models that perform well in sample images (usually the last 2-3 rounds of models) for saving.

Select and test them in the Model section of your personal center.

Select your own LoRA in the creation flow for image generation testing.

<figure><img src="/files/q5TmtPyIcY7hp1X5W4tq" alt=""><figcaption></figcaption></figure>


# Flux Lora Training

Learn to train Flux-related Lora models online in three steps.

**Page Entry:** Click the <mark style="background-color:yellow;">"Train"</mark> button in the top left to enter Lora Training.

## Three Steps to Train Flux Lora

### STEP1: Create a Dataset

We recommend using at least <mark style="background-color:yellow;">**30**</mark> images.

### STEP2: Set the Parameters

<mark style="color:red;">**Note:**</mark> Choose <mark style="background-color:yellow;">**Flux**</mark> as the base model.

For details on other parameter settings, refer to:

{% content-ref url="/pages/bDLwl9DD9y2ga18oYqz4" %}
[3-2  LoRA Training (Advance)](/guide-1/3-advanced-guide/3-2-lora-training-advance)
{% endcontent-ref %}

### STEP3: Click " Train Now"

After the training is complete, select a Lora and click <mark style="background-color:yellow;">"Publish."</mark>


# Edit Model Training

## kontext Training

## Differences between Kontext training and conventional LoRA training:

| Comparison   | Training Principle                                                        | Number of Training Images                    | Applicable Scenarios                                                                            | Generalization Ability                                      |
| ------------ | ------------------------------------------------------------------------- | -------------------------------------------- | ----------------------------------------------------------------------------------------------- | ----------------------------------------------------------- |
| LoRA         | Traditional low-rank adaptation training, mainly learning visual features | Generally requires 50-100 images or more     | Style transfer, character/object feature learning                                               | Strong generalization ability with similar visual features  |
| Kontext LoRA | Context-based training method, focusing on semantic associations          | Usually 20-50 pairs can achieve good results | Creative scenarios requiring precise semantic control, image modification, and partial transfer | Strong generalization ability in similar semantic scenarios |

**Basic Concepts and Architecture**

\- Difference from Flux Dev:

Flux Dev is Text to Image (prompt → target image), Kontext is image conversion (prompt + original image → target image).

\- Application Scenarios: Image editing, style conversion, content modification, etc.

<figure><img src="/files/oojpABvhcPyUGeOGch9x" alt=""><figcaption></figcaption></figure>

Flux LoRA

Kontext LoRA

Flux.Kontext and Kontext.Fast differ only in training speed, and the consumption is relatively higher for Kontext.Fast than Flux.Kontext

#### Dataset Preparation

**Image Screening**

● Kontext training dataset images must be in pairs: each group includes one original image, one result image, and annotations.

● Number of image sets: 20-50 groups recommended, maximum upload limit: 100 groups.

● Image selection: Multiple types of original images, watermarked images need watermark removal.

Example: For conversion to ink style, original images can include realistic style, line art style, people, objects, and landscapes to increase generalization.

**Image Annotation**

● Annotation language: English annotations are recommended to avoid translation errors.

● Default annotation: Can be filled in when no annotation is available. Not needed if all groups are annotated.

● Annotation Content:

Correct annotations: Describe the difference between the original and the result images with simple descriptions.

Correct examples: Convert image to ink style; Transform into fisheye lens; Turn character into a cat.

Incorrect annotations: Like regular image model training annotations, which describe the image content.

Incorrect example: A slim young Asian woman with long black hair, wearing a pink tank top and blue denim shorts, standing on a green railing by the river with a bridge in the background.

#### Dataset Upload

**Dataset Upload**

● Dataset naming format: "filename\_start.png", "filename\_end.png", "filename.txt".

● If not named in a fixed format, manual matching of images and re-annotation is required.

**Image Upload**

● Upload all images and fill each group in order. If the original and result images don't match correctly, manually move the images to the corresponding groups.

Flux.Kontext Upload Dataset

l  The dataset naming format: ctrl0, target

l  ctrl0, refers to the folder where the original image is placed. target, refers to the folder where the result image and annotation are placed.

l  Note: You need to compress the ctrl0 and target folders together into a dataset package for successful upload.

l  Original images in "ctrl0" should be named as: 1\_0.png, 2\_0.png, 3\_0.png

l  Result images and annotations in "target" should be named as: 1.png, 1.txt, 2.png, 2.txt, 3.png, 3.txt

Note: Regarding the dataset upload method, Flux.Kontext uses the same format as qwen image edit for uploading datasets; both require uploading compressed packages.

Image Upload

l  Click to upload images. Up to 50 images can be uploaded at once. Those beyond this limit cannot be uploaded successfully.

l  Select all images to upload. If there is no fixed naming format, they will be placed in the unmatched images section directly below the matching group. Up to 50 unmatched images are allowed, and you can manually move them to the corresponding groups.

#### Parameter Setting

**Training Steps**

● Training steps optimization: The model needs to show proper effects across all images, maintaining better generalization. This requires more dataset images and increased steps.

For limited specific use, choose fewer dataset images and reduce steps.

● Recommended steps: Default is 1,500 steps, for a dataset of 20-50 groups, 2,000-5,000 steps are recommended.

**Training Epoch**

l  Times per Image Repeat: Calculated based on the number of groups, indicating training times per group.

l  Cycles Epoch: When all groups complete training according to the Times per Image Repeat count, one Cycles Epoch is completed. This is consistent with the Epoch and Repeat concepts in image training.

l  In most cases, using the default parameters for training is sufficient. Generally, when there are more groups, you can select fewer Times per Image Repeat and Cycles Epoch.

**Learning Rate**

● For a dataset of 20-50 groups, the default learning rate should be fine. If there are more image groups and steps, and you want the model to be effective consistently, increase the learning rate.

**Default Annotation**

● If all or any image groups lack annotations, the default annotation will be used to avoid missing annotations.

● Example 1:

<figure><img src="/files/nQxtLrAD65gscGYbrxhH" alt=""><figcaption></figcaption></figure>

Change the photo of green stalks with yellow bananas to black and white line drawings.

● Example 2:

Change the woman in the picture into a cat, keeping the background the same, the clothing and decorations the same.

#### Model Testing

**Model Saving**

● After training completion, only one model is produced. Click to save the model.

**Testing Method**

● We recommend testing via a workflow.

[SeaArt AI | Kontext](https://www.seaart.ai/workFlowAppDetail/d1rqh8u6sm8c73e739sg)

<figure><img src="/files/82hK0e9p2YDfG9MkfDU8" alt=""><figcaption></figcaption></figure>

● Test Parameter Settings:

Please select the LoRA name: Select the saved model.

Please enter the model strength: Generally, use the default value 1.

Please enter the width: The output image's width.

Please enter the height: The output image's height.

Please select an image: Select the image to be modified.

Please enter text: The prompt, a short sentence similar to or identical to the annotation content.

### Qwen-Image-Edit Training

| Comparison           | Base Model            | Core Feature                    | Multi-image Editing |
| -------------------- | --------------------- | ------------------------------- | ------------------- |
| Kontext              | FLUX.1                | Context-Aware Image Generation  | 1                   |
| Qwen-Image-Edit      | Qwen-Image            | Image Editing and Understanding | 1                   |
| Qwen-Image-Edit-2509 | Qwen-Image (Enhanced) | Enhanced Image Editing          | 1-3                 |

#### Dataset Preparation

Image Screening

l  Qwen-Image-Edit training dataset images must be in pairs: each group includes one original image, one result image, and one annotation.

l  Number of image sets: 10-30 groups recommended, maximum upload limit: 100 groups.

l  Image selection: Multiple types of original images, watermarked images need watermark removal.

Example: For conversion to ink style, original images can include realistic style, line art style, people, objects, and landscapes to increase generalization.

Image Annotation

l  Annotation language: English annotations are recommended to avoid translation errors.

l  Default annotation: Can be filled in when no annotation is available. Not needed if all groups are annotated.

l  Annotation content:

Correct annotations should describe the differences between the original and the result images with simple descriptions.

Correct examples: Convert image to ink style; Transform into fisheye lens; Turn character into a cat.

Describing the image content like regular image model training annotations is incorrect.

Incorrect example: A slim young Asian woman with long black hair, wearing a pink tank top and blue denim shorts, standing on a green railing by the river with a bridge in the background.

l  The image annotation is consistent with Kontext dataset annotation, with the only difference being that Qwen-Image-Edit supports Chinese annotation.

#### Dataset Upload

Qwen-Image-Edit Single Image Dataset Upload Method

l  The dataset naming format: ctrl0, target

l  ctrl0, refers to the folder where the original image is placed. target, refers to the folder where the result image and annotation are placed.

l  Note: You need to compress the ctrl0 and target folders together into a dataset package for successful upload.

l  Original images in "ctrl0" should be named as: 1\_0.png, 2\_0.png, 3\_0.png

l  Result images and annotations in "target" should be named as: 1.png, 1.txt, 2.png, 2.txt, 3.png, 3.txt

Qwen-Image-Edit 2509 Dual Image Dataset Upload Method

l  The dataset naming format: ctrl0, ctrl1, target

l  ctrl0, refers to the folder where the original image is placed.ctrl1, refers to the folder where the original image 2 is placed. target, refers to the folder where the result image and annotation are placed.

l  Note: You need to compress the ctrl0, ctrl1, and target folders together into a dataset package for successful upload.

![descript](/files/vuSv9lSA6kt5iLz9YlDm)

l  Original images in "ctrl0" should be named as: 1\_0.png, 2\_0.png, 3\_0.png

![descript](/files/pOS0Rla6cOVPn3YCo10Z)

l  Original images in "ctrl1" should be named as: 1\_1.png, 2\_1.png, 3\_1.png

![descript](/files/R84cewX24tOG1KAjRret)

l  Result images and annotations in "target" should be named as: 1.png, 1.txt, 2.png, 2.txt, 3.png, 3.txt

![descript](/files/fLu25oSa2rBIyJtEn7As)

l  First, set the group picture quantity to dual image training.

Under the dual image training group, select a compressed package containing the ctrl0, ctrl1, and target folders.

This will complete the dataset upload.

![descript](/files/UCPZex2VpoZncGlo1oQ1)

Qwen-Image-Edit 2509 Three Image Dataset Upload Method

l  The dataset naming format: ctrl0, ctrl1, ctrl2, target.

l  ctrl0, refers to the folder where the original image is placed.ctrl1, refers to the folder where the original image 2 is placed.ctrl2, refers to the folder where the original image 3 is placed. target, refers to the folder where the result image and annotation are placed.

l  Note: You need to compress the ctrl0, ctrl1, ctrl2, and target folders together into a dataset package for successful upload.

![descript](/files/6ZOKdUAkLa7LtwJPmHMF)

l  Original images in "ctrl0" should be named as: 1\_0.png, 2\_0.png, 3\_0.png

l  Original images in "ctrl1" should be named as: 1\_1.png, 2\_1.png, 3\_1.png

![descript](/files/GFEoKlPeZ6BFigINnkx6)

l  Original images in "ctrl2" should be named as: 1\_2.png, 2\_2.png, 3\_2.png

![descript](/files/vkv5YM2i8JqYLfdhiONW)

l  Result images and annotations in "target" should be named as: 1.png, 1.txt, 2.png, 2.txt, 3.png, 3.txt

![descript](/files/AotTwhwopt2nkjNNJvDR)

l  If not named in a fixed format, manual matching of images and re-annotation is required.

Image Upload

l  Click to upload images. Up to 50 images can be uploaded at once. Those beyond this limit cannot be uploaded successfully.

l  Select all images to upload. If there is no fixed naming format, they will be placed in the unmatched images section directly below the matching group. Up to 50 unmatched images are allowed, and you can manually move them to the corresponding groups.

Parameter Settings

Training Epoch

l  Times per Image Repeat: Calculated based on the number of groups, indicating training times per group.

l  Cycles Epoch: When all groups complete training according to the Times per Image Repeat count, one Cycles Epoch is completed. This is consistent with the Epoch and Repeat concepts in image training.

l  In most cases, using the default parameters for training is sufficient. Generally, when there are more groups, you can select fewer Times per Image Repeat and Cycles Epoch.

Learning Rate

l  For a dataset of 20-50 groups, the default learning rate should be fine. If there are more image groups and steps, and you want the model to be effective consistently, increase the learning rate.

Default Annotation

l  If all or any image groups lack annotations, the default annotation will be used to avoid missing annotations.

l  For example, if the annotation for Group 3 is missing, but we have set a default annotation, then it will be automatically added to Group 3 when training begins.

Model Effect Preview Original Image

For preview images, use the original images (upload as many original images as you have).

Model Effect Preview Annotation

The preview annotation needs to be consistent with the dataset annotation.

Complete all settings and click "Start Training" to begin (The training methods for Qwen Image Edit, Qwen Image Edit 2509, and Flux.kontext are basically the same).

Model Testing

Model Saving

l  The number of saved models is calculated based on epochs. There will be as many models saved as there are epochs. Select the one with the best sample image effect, and proceed with testing once the model synchronization is complete.

Testing Method

l  We recommend testing via a workflow.

l  Qwen Image Edit LoRA Test Workflow

[SeaArt AI | Qwen Edit](https://www.seaart.ai/workFlowAppDetail/d48rh6te878c73essh90)

l  Test Parameter Settings

Please select an image: Input the image that requires editing.

Please select the LoRA name: Select the trained LoRA.

Please enter the model strength: LoRA's weight, which usually doesn't need adjustment.

Please enter prompt: Enter the prompts (you can directly enter prompts from training tags).

l  Qwen Image Edit 2509 LoRA Test Workflow (Single Image Editing Test)

[SeaArt AI AI | Qwen Edit 2509](https://www.seaart.ai/zhCN/workFlowAppDetail/d49eh6de878c73cflk6g)

l  Test Parameter Settings

Please select an image: Input the image that requires editing.

Please select the LoRA name: Select the trained LoRA.

Please enter the model strength: LoRA's weight, which usually doesn't need adjustment.

Please enter prompt: Enter the prompts (you can directly enter prompts from training tags).

l  Qwen Image Edit 2509 LoRA Test Workflow (Dual Image Editing Test)

[SeaArt AI | Dual Image Editor Qwen Edit 2509](https://www.seaart.ai/zhCN/workFlowAppDetail/d48smnle878c73d8qrjg)

![descript](/files/lytxlDyKIlgKH0U5CeUH)

l  Test Parameter Settings

Please select an image: Input the image that requires editing (the two images need to maintain type consistency with the dataset content).

Please select the LoRA name: Select the trained LoRA.

Please enter the model strength: LoRA's weight, which usually doesn't need adjustment.

Please enter prompt: Enter the prompts (you can directly enter prompts from training tags).

l  Qwen Image Edit 2509 LoRA Test Workflow (Three Image Editing Test)

[SeaArt AI | Three Image Editor Qwen Edit 2509](https://www.seaart.ai/workFlowAppDetail/d48s8vle878c73fmdab0)

![descript](/files/YvppSvycI2SGhunrhBLk)

l  Test Parameter Settings

Please select an image: Input the image that requires editing (the three images need to maintain type consistency with the dataset content).

Please select the LoRA name: Select the trained LoRA.

Please enter the model strength: LoRA's weight, which usually doesn't need adjustment.

Please enter prompt: Enter the prompts (you can directly enter prompts from training tags).


# Video Training

## Preprocessing Videos

### Selection of Training Videos

* Use videos with consistent content, actions, or visual effects, but different main subjects.
* Prioritize using videos; images can be used as supplementary data.
* Videos must be high-resolution and watermark-free.

### **Number of Videos**

* 4 to 10 videos are sufficient. (Image-only training is not recommended.)

### **Frame Rate**

* Convert videos to 16fps, with a total of 81 frames (i.e., 5 seconds in duration).
* You can use video editing tools to trim clips to 5 seconds, then extract frames at 16fps.
* Shorter videos (e.g., 2s or 3s) are also acceptable, but they **must** be processed to 16fps.

### **Resolution**

* 480p works well. You can also reduce it to 320p to speed up training.
* (Training will likely fail if the resolution is too high.)

### Video Tagging

Automatic Tagging

Manual Tagging

Key Points: Secondary Features + Main Features

Main Features: Actions/effects to be learned; Secondary Features: Characters in the video, where they are, what they're doing.

Example: In the video, a woman wearing a black formal suit is presented. The person raises her hand and showers colorful confetti in celebration with a smile. The person then reveals a bikini, causing a b1k1n1 bikini up effect. The person continues celebrating, further showing the b1k1n1 bikini up effect.

The part before the red text describes the video content, and the red text summarizes the action effects being learned (i.e., red text represents main features, the rest are secondary features).

[3(2).mp4](https://drive.weixin.qq.com/s?k=AFUA3QfrAA83nXz9wHAXUA2wbXAPE)

### Online Training

#### Video Model Introduction

Hunyuan Video

Text-to-video: hunyuanvideo-fp8

Wan Video

Text-to-video: Wan2.1-14B

Image-to-video: Wan2.1-14B-480P, Wan2.1-14B-720P

Difference between text-to-video and image-to-video: In parameter adjustment, for model effect preview prompts, text-to-video only needs text similar to training set tags to generate preview images.

Image-to-video requires inputting images and corresponding prompts to generate preview images.

### Wan 2.1 Video LoRA Training

#### Video Model Introduction

Wan Video

Text-to-Video: Wan2.1-14B.

Image-to-Video: Wan2.1-14B-480P, Wan2.1-14B-720P.

Difference between Text-to-Video and Image-to-Video: In parameter settings, under Model Effect Preview Prompts, for text-to-video, you only need to enter text similar to the training set captions to generate preview samples.

For image-to-video, you must provide both an image and the corresponding prompt to generate preview samples.

###

### Online Parameter Settings

#### Image-to-video

Image-to-video: Wan2.1-14B-480P, Wan2.1-14B-720P (mainly selected based on training video resolution).

For training materials of 216\*320 (less than 480p), choose the 480p model (there is little difference in final training effect between 720p and 480p, so 480p is recommended).

| Resolution | Specific Size | Total Pixels  |
| ---------- | ------------- | ------------- |
| 480p       | 854\*480      | About 410,000 |
| 720p       | 1280\*720     | About 920,000 |

Complete Dataset Upload

#### Parameter Settings

Frames to Extract: Number of frames to extract from a single video segment.

Example: For each segment at 16fps, setting Frames to Extract to 9 means not every frame will be learned.

Number of Slices: Dividing each video material.

Example: For a 5-second video at 16fps, setting Number of Slices to 5 means each segment is 16 frames; if set to 4, each segment is 20 frames.

Times per Image: Learning times for each video.

Cycles: Number of cycles based on Times per Image.

Model Effect Preview Prompts: Prompt for generating example video (modify it based on dataset tags combined with initial frame image content).

Initial frame: For image-to-video, the required image for generating the example video.

#### Advanced Parameter Settings

The only setting to be modified: Flow Shift.

720p is 5, 480p is 3 \[materials must also be 480p].

<figure><img src="/files/OaDO7DwVyeVSv4lLWte7" alt=""><figcaption></figcaption></figure>

#### Text-to-Video Parameters

Text-to-video parameters are consistent with image-to-video parameters. Flow Shift follows the default parameters.

#### Model Selection

Choose the one with good real-time sample images that match the effects or actions shown in the training set videos.

### Model Testing

#### Image-to-Video Testing

kijai Workflow: [kj wan testing.json](https://drive.weixin.qq.com/s?k=AFUA3QfrAA8Q0yaG1PAXUA2wbXAPE)

AI App Testing: [SeaArt AI AI | kj wan testing](https://www.seaart.ai/zhCN/workFlowAppDetail/d1e017de878c73eofvrg)

Parameter Settings

Model Selection: The training model should match the testing model.

Select LoRA: Choose saved LoRA from your models.

Weight: LoRA weight.

Width: The size after the input image is compressed and cropped.

Height: The size after the input image is compressed and cropped.

Frames: Total frame count for the output duration (calculated as 4\*n+1, where n represents seconds; e.g., 5 seconds = 81 frames).

Shift: 720p is 5, 480p is 3.

CFG: Default cfg is 6, can be adjusted to 5.

Official Workflow: [wan official workflow.json](https://drive.weixin.qq.com/s?k=AFUA3QfrAA8xS1OCycAXUA2wbXAPE)

AI App Testing: [SeaArt AI AI | wan official workflow](https://www.seaart.ai/zhCN/workFlowAppDetail/d1e0fvte878c73bq8t80)

Default cfg is 6, can be adjusted to 5.

Sampler is uni-pc, scheduler can be normal or simple.

Sampler dpmpp\_2m, scheduler sgm\_uniform.

Note: Other parameters are consistent with the kj parameter settings.

#### Text-to-Video Testing

Wan Creation Flow Testing

Model: Select wan2.1.

Additional: Select saved trained model.

Select Text to Video.

Hunyuan Creation Flow Testing

Model: Hunyuan Video.

Additional: Select saved trained model.

Select Text to Video.

### Hunyuan LoRA Video Training

#### Video Model Introduction

Hunyuan Video: Currently, only online text-to-video training is available.

Text-to-Video: hunyuanvideo-fp8.

#### Parameter Settings

Frames to Extract: Number of frames to extract from a single video segment.

Example: For each segment at 16fps, setting Frames to Extract to 9 means not every frame will be learned.

Number of Slices: Dividing each video material.

Example: For a 5-second video at 16fps, setting Number of Slices to 5 means each segment is 16 frames; if set to 4, each segment is 20 frames.

Times per Image Repeat: Learning times for each video.

Cycles Epoch: Number of cycles based on Times per Image Repeat.

Model Effect Preview Prompts: Prompt for generating example video (modify it based on dataset tags combined with initial frame image content).

<figure><img src="/files/OkIarpmBNPeyFo6sbdYc" alt=""><figcaption></figcaption></figure>

Hunyuan Creation Flow Testing

Model: Hunyuan Video.

Additional: Select saved trained model.

Select Text to Video.

### Wan 2.2 Video LoRA Training

Video preprocessing is the same as Wan 2.1: resolution, video length, video fps, and video quantity are the same.

#### Video Model Introduction

Wan Video

Text-to-Video: wan2.2 t2v-low, wan2.2 t2v-high.

Image-to-Video: wan2.2 i2v-low, wan2.2 i2v-high.

Differences between Wan 2.2 Video and Wan 2.1 Video Training:  Text-to-video and image-to-video each have two models: a high-noise model and a low-noise model.\
High-noise models mainly control motion/dynamics in the video, while low-noise models mainly control fine details.\
To minimize training time, you can only train the low-noise model to reach a basic effect quickly.

It is best to train both high-noise and low-noise models on the same dataset, and then load two LoRAs together. This can fully utilize the strong capabilities of Wan 2.2.

The Wan 2.2 video models have stronger language understanding, so you can use simpler and more unified language descriptions when annotating videos.

Wan 2.2 Video Annotation

Wan 2.2 Text-to-Video

Automatic Annotation

You can use Automatic Annotation.

Manual Annotation

Describe the video content clearly; avoid overly generic descriptions such as "a person," "an animal."

#### Wan 2.2 Image-to-Video

Automatic Annotation

Automatic annotation is not recommended, as there are only a few video assets, and it tends to be overly detailed. For Wan 2.2, excessively detailed captions can make it more cumbersome to enter prompts when using the LoRA later.

Manual Annotation

You can use simple descriptions if all video assets are of a person.

For example: A person, whose head turns into a pumpkin head, then wears a robe, and a pumpkin lantern, bats, and a moon Halloween background.

You do not need to specify gender or age in the description; simply describe “a person,” clearly and consistently outlining the transformation effects. Then, copy this entire sentence into the captions of other videos.

#### Parameter Setting

Frames to Extract: Number of frames to extract from a single video segment.

Example: For each segment at 16fps, setting Frames to Extract to 9 means not every frame will be learned.

Number of Slices: Dividing each video material.

Example: For a 5-second video at 16fps, setting Number of Slices to 5 means each segment is 16 frames; if set to 4, each segment is 20 frames.

Times per Image Repeat: Learning times for each video.

Cycles Epoch: Number of cycles based on Times per Image Repeat.

Model Effect Preview Prompts: Prompt for generating example video (modify it based on dataset tags combined with initial frame image content).

Initial frame: For image-to-video, the required image for generating the example video.

#### High/Low Noise Model Effect Comparison

The table shows the impact of low-noise and high-noise LoRA on the final generated videos.

From this we can see that even if you only train a high-noise model, the overall video effect can basically be expressed. But some fine details still need the low-noise model to present them.

#### Image to Video

Model: Generally, just choose wan2.2 i2v-high.

Advanced Parameter Settings: The default values are the best.

**Model Effect Preview Prompts: You can directly use the video captions.**

Initial Frame: It is recommended that the uploaded image type be consistent with the video asset type. For example, if the video asset is realistic, upload a realistic image; if the video is a half-body shot, upload a half-body image.

#### Text to Video

Model: Choose wan2.2 t2v-low.

Model Effect Preview Prompts: You can directly use the video captions.

#### Model Testing

Generally, save and test the LoRAs from the last few epochs.

#### Image-to-Video Testing

AI App Testing: [SeaArt AI AI | WAN 2.2 Test](https://www.seaart.ai/zhCN/workFlowAppDetail/d4ihmble878c738mrk2g)

<figure><img src="/files/wpCepcx6htK3n6Nfmx4d" alt=""><figcaption></figcaption></figure>

**Parameter Settings**

First LoRA: Select the high-noise LoRA.

Second LoRA: Select the low-noise LoRA.

If you only have a high-noise LoRA, then choose the same high-noise LoRA for the low-noise slot and set its strength to 0. If you have a low-noise LoRA, you can use both high-noise and low-noise LoRAs together.

Please enter text: Input the prompts (captions).

Please select an image: Input the image (preferably similar to the first frame of the training video).

#### Text-to-Video Testing

AI App Testing: [SeaArt AI | wan2.2 t2i Test](https://www.seaart.ai/workFlowAppDetail/d4ruebte878c73dv1lsg)

Parameter Settings

First LoRA: Select the high-noise LoRA.

Second LoRA: Select the low-noise LoRA.

If you only have a high-noise LoRA, then choose the same high-noise LoRA for the low-noise slot and set its strength to 0. If you have a low-noise LoRA, you can use both high-noise and low-noise LoRAs together.

Please enter text: Input the prompts (captions).

Please select an image: Input the image (preferably similar to the first frame of the training video).

&#x20;


# 3-3 Workflow Guide

Unlock ComfyUI's image conversion power! This simple guide explores advanced techniques for precise style modifications, facial referencing, and high-definition results.

If you want to dive deeper into Workflow, our [ComfyUI WIKI](https://docs.seaart.ai/seaart-comfyui-wiki) should be your go-to station.


# Image Conversion

Master ComfyUI's image conversion! This guide enhances Image-To-Image with precise control, facial referencing, and high-definition upscaling.

**Thinking Process:**

Comfyui's Image conversion is similar to webui's Img2Img, where an original image is uploaded and its style is modified through the model. However, to enhance the precision of the conversion, we can add a few new steps:

> **I. Add a model magnification node to control the size of the original image.**&#x20;
>
> **II. Increase similarity to the original image:**&#x20;
>
> a. Use ipadapter faceid to reference facial features.&#x20;
>
> b. Reverse-engineer the original image prompts.&#x20;
>
> c. Add ControlNet (openpose, canny, depth).&#x20;
>
> **III. Upscale the final Image**

<figure><img src="/files/IdplIjLVeQQ2hy7s6o3o" alt="ComfyUI Image Conversion Process"><figcaption></figcaption></figure>

### Step 1: Build the Model Group

We can start building on the basis of the Img2Img template. First, selectively add an "Upscale Image By" node behind the original image to control the size of the original image. Lora can be added according to individual needs, or not added at all. Add a "CLIP Set Last Layer" node according to your needs, which can also be omitted. This node allows skipping of layers, and finally, connect the corresponding nodes.

**Add nodes:**

Upscale Image By&#x20;

CLIP Set Last Layer

### Step 2: Reference the Original Image

<mark style="background-color:yellow;">**(Reverse-engineer prompts + Ipadapter + Control Net)**</mark>

1. Reverse-engineer prompts: **WD14 Tagger node**&#x20;

Double-click to search and add the WD14 Tagger node.&#x20;

Connect the image node.&#x20;

Right-click on the positive prompts node and select "**Convert text to input**" to connect the WD14 Tagger to the positive prompts node.

<figure><img src="/files/QLqSCTKElKt0JR1A26If" alt="ComfyUI Image Conversion - BReverse-engineer Prompts: WD14 Tagger Node 1" width="563"><figcaption></figcaption></figure>

However, this only includes the prompts from the image. If you want to add other prompts, You need to create a new **Text Concatenate** node, which can connect multiple segments of prompts together.&#x20;

Then create a new **Primitive** node. The Primitive node can be connected to any node to become a related attribute.&#x20;

Enter additional prompts on the Primitive, such as Lora Trigger Words, some quality words, etc.

<figure><img src="/files/f322zx4X5l9smSgN40oJ" alt="ComfyUI Image Conversion - BReverse-engineer Prompts: WD14 Tagger Node 2" width="362"><figcaption></figcaption></figure>

<figure><img src="/files/6YbI0d0Q7P4DeXzzcxxf" alt="ComfyUI Image Conversion - BReverse-engineer Prompts: WD14 Tagger Node 3" width="563"><figcaption></figcaption></figure>

At this point, the prompts not only include the ones reverse-engineered from the image but also those we input.

2. Next, set up the **IPadapter FaceID** to reference facial features:&#x20;

Double-click to search for IPadapter FaceID and match the input nodes accordingly.

<figure><img src="/files/Ta1Lsz1QSJQOHpdszDAk" alt="ComfyUI Image Conversion - Set up the IPadapter FaceID to Reference Facial Features"><figcaption></figcaption></figure>

After dragging out the node, create new nodes:

ipadapter→ IPAdapter Model Loader&#x20;

clip\_vision→ Load CLIP Vision&#x20;

insightface→ IPAdapter InsightFace Loader

Connect the output to the sampler.

**Add nodes:**

IPadapter FaceID

IPAdapter Model Loader

Load CLIP Vision

IPAdapter InsightFace Loader

3. **Set up the ControlNet**&#x20;

It is recommended to use the **CR Multi-ControlNet Stack** node, which allows the addition of multiple ControlNets. Then add the corresponding preprocessors. It is recommended to use **OpenPose, Canny, and Depth** as the ControlNet. You can add or remove them based on the final visual needs. Afterwards, add a ControlNet application node at the output: **CR Apply Multi-ControlNet**. It is recommended to set the resolution at the preprocessor to **1024**.&#x20;

<mark style="background-color:yellow;">\*Finally, remember to turn on the switches of the ControlNet that will be used.</mark>

<figure><img src="/files/Cccx7PI3VQ1HuI83myfu" alt="ComfyUI Image Conversion - Set up the ControlNet" width="563"><figcaption></figcaption></figure>

**Input:**&#x20;

Connected to the positive and negative prompt nodes.

&#x20;**Output:**&#x20;

Connected to the sampler.

**Add nodes:**

CR Multi-ControlNet Stack

CR Apply Multi-ControlNet

### Step 3: High-Definition Restoration

After the model group and reference to the original image have been set up, we can add an image high-definition restoration step at the final output of the image:

<figure><img src="/files/37J6bQaHd0RPYXo1GK1P" alt="ComfyUI Image Conversion - High-Definition Restoration" width="562"><figcaption></figcaption></figure>

**Add Nodes:** Upscale Image (using Model)

After assembling the nodes, you can organize them into a group for easier viewing.

<figure><img src="/files/SCO6wUQjLPeXdyfI0W9D" alt="ComfyUI Image Conversion - Add Nodes to Upscale Image 2"><figcaption></figcaption></figure>

Finally, it's necessary to adjust the relevant parameters based on the output image, <mark style="background-color:yellow;">such as ckpt, Lora weights, prompt words, sampler, redraw scale, etc.</mark>&#x20;

**Key parameters for this conversion include:**&#x20;

CLIP \_layer：-2&#x20;

Upscale Image By: 1.5&#x20;

steps: 40&#x20;

sampler\_name: dpmpp\_2m&#x20;

scheduler: karras&#x20;

denoise: 0.7

<mark style="color:red;">Note: If you choose the SDXL model, you will also need the corresponding SDXL Lora, and adjust the ControlNet to SDXL; otherwise, the image output will fail.</mark>

<figure><img src="/files/KXV0PncM44kRIrB7tZq2" alt="ComfyUI Image Conversion - Corresponding SDXL Lora" width="239"><figcaption></figcaption></figure>

The above is a complete workflow for image conversion. Based on this, you can also add VAE or FreeU\_V2 to adjust the final image:&#x20;

**FreeU\_V2:** Mainly controls color and extracts some content for optimization.&#x20;

**Load VAE:** Fine-tunes the color and details of the image.

Through such a workflow, you can achieve conversions to different styles.


# Inpainting

Learn inpainting and modifying images in ComfyUI! This guide covers hair, clothing, features, and more using "segment anything" for precise edits.

### **Change hair color, hairstyle, chest, abs, clothing, etc.**

Before starting to build nodes, it's helpful to establish a workflow.

For example, if you want to change hair color, you'll need these two steps:

**Step 1:** First, the hair needs to be identified. You can choose manual painting or automatic recognition. **Step 2:** Partial repaint, modify the recognized area.

<figure><img src="/files/8NA7MMxX8mtYFrZoIM6Z" alt=" ComfyUI Redraw - Change Hair Color Process" width="504"><figcaption></figcaption></figure>

Now, let's start building nodes based on these steps.

<mark style="background-color:red;">**Method 1: segment anything**</mark>

**Step 1: Identify the hair**

I. After uploading the image, add a 'segment anything' node. Drag the node out and load two corresponding models. Connect the image to the node, input the desired area for recognition.

> **Add nodes:**&#x20;
>
> **GroundingDinoSAMSegment (segment anything)**：SAMLoader (Impact)、GroundingDinoModelLoader (segment anything)&#x20;
>
> **Parameters:**&#x20;
>
> device\_model: Prefer GPU

II. To enhance the accuracy of mask recognition, add another 'segment anything' node to identify the face. You can use the same model as before for input. Finally, subtract the two masks using 'mask-mask': mask1 - mask2 to obtain precise hair parts.

**Add nodes:**

**Bitwise(MASK - MASK)**

<mark style="background-color:yellow;">\*When encountering areas that cannot be recognized, you can use this method of subtracting masks.</mark>

III. After obtaining the mask, you can perform some processing on the mask. For example, here we added 'GrowMask' and 'feathered mask' to expand the mask and feather the edges, making the final inpainting image more natural.

**Add nodes:**

**GrowMask**

**FeatheredMask**

IV. Finally, you can add a mask preview node to view the final mask effect.

<figure><img src="/files/mVSk2jDuVIBPfT8sqbgd" alt="ComfyUI Redraw - Segment Anything - Mask Preview Node" width="422"><figcaption></figcaption></figure>

**Add nodes:**

**Convert Mask to Image**

**Preview Image**

V. For easier viewing, we can organize this group of nodes together.

**Step 2: Partial Repaint**

I. Here, for the inpainting, we'll use the previously mentioned 'Set Latent Noise Mask', which will reference the original image for the inpainting. Connect the mask nodes together.

<figure><img src="/files/WLDKM8OnKFSr2oaLM1i1" alt="ComfyUI Redraw - Partial Repaint - Set Latent Noise Mask" width="513"><figcaption></figcaption></figure>

**Add nodes:**

**Set Latent Noise Mask**

II. Next, add the model group. Follow the Img2Img, add models prompts, sampler, and finally VAE Decode. Then connect the mask group to the model group.

<mark style="background-color:yellow;">\*The images need to be encoded to enter the latent space for inpainting</mark>

III. All lines are connected. You can input corresponding prompts according to your needs, set relevant parameters, and the maximum Denoising Strength is 1.

Note that when inpainting the face, it's advisable to add the FaceDetailer node, which helps to enhance facial details. Connect the input end to the corresponding node.

**Add nodes:**

**FaceDetailer:** UltralyticsDetectorProvider、SAMLoader (Impact)

<figure><img src="/files/vK59tWUY5u0DIRYypoLH" alt="ComfyUI Redraw - FaceDetailer"><figcaption></figcaption></figure>

Additionally, since we only need to repaint the face and don't need to subtract two masks, you can select the corresponding node, press Ctrl+B to hide the node, and then connect the corresponding nodes.

This type of inpainting will redraw the entire face, essentially making it look like a different person. Regarding how to achieve different expressions for the same person, a more detailed tutorial will be released later. Finally, we can save this workflow for future use.

Using this method, we can achieve many one-click inpaintings.

Such as One-click Breast Enhancement , one-click muscle gain, one-click facial modification, one-click change of clothing, hairstyle, etc.

<mark style="background-color:red;">**Method 2: Yoloworldwodel ESAM**</mark>

Capable of quickly extracting a specific object, with relatively less precise recognition compared to segmen anything. Some parts (such as the belly) may not be detectable. It can be used in conjunction with mask-mask.

I. First, follow the same steps as "segment anything" and integrate the Yoloworld ESAM automatic detection feature to identify areas that need to be redrawn. Taking 'One-click muscle gain' as an example, since it's unable to directly recognize the areas needing muscle enhancement, the mask-mask approach is used to obtain the regions for enhancement. Then, process the masked areas accordingly.

<mark style="background-color:yellow;">\*You can upscale the image on the original image before starting.</mark>

> **Important parameters:**&#x20;
>
> **confidence\_threshold:** Accuracy of recognition. A smaller value results in more precise recognition. **iou\_threshold:** Degree of overlap for bounding boxes. A smaller value results in more precise recognition.&#x20;
>
> **mask\_combined:** Whether to overlay masks. If "true," the mask will be combined and output on a single image. If "false," the mask will be output separately.

II. Integrate the mask into the "InpaintModelConditioning" node for inpainting masks. The prompts are only effective for masked areas. According to theImg2Img mode, add the model-prompt-sampler-decode-preview image.

<figure><img src="/files/Tpuc56fYj8N04vUJ5AGG" alt="ComfyUI Redraw - InpaintModelConditioning"><figcaption></figcaption></figure>

III. Adjust relevant parameters, then click "Generate.”

Here are more [examples of ComfyUI Workflow](https://docs.seaart.ai/seaart-comfyui-wiki/comfyui-workflow-example) for your inspiration.


# 3-4 Canvas Guide

SeaArt Canvas: Design stunning e-commerce posters with AI! This guide explores real-time editing, background generation, and seamless product integration.

> SeaArt Canvas is divided into Realtime Canvas and Generate Canvas.&#x20;
>
> In **Realtime Canvas**, we can view the image results in real-time and make adjustments to address the uncontrollable issues of AI image generation.&#x20;
>
> In **Generate Canvas**, we can directly use the image generation function on the canvas, enabling convenient use AI functions such as partial repaint and erasure on images.

SeaArt Canvas offers deep involvement in the design process, with highly controllable AI drawing capabilities that turn your imagination into reality. Next, let's experience the functionality of SeaArt Canvas through a practical e-commerce poster example.

## **Product Background**

<figure><img src="/files/vvuAXzaRn5hdBHVomC9B" alt="Product Background" width="188"><figcaption></figcaption></figure>

【Brand New Upgrade · Ultimate Experience】The XS4 is a meticulously crafted gaming controller, utilizing the latest technology and premium materials to bring you an unprecedented gaming experience. Its ergonomic design ensures comfortable grip even during extended gaming sessions, while responsive buttons and precise joysticks guarantee accurate operation with every input.

&#x20;【Wide Compatibility】Whether it's PC, PS4, Xbox, or Switch, this multi-platform gaming controller is perfectly compatible. With no complicated setup required, it's plug-and-play, allowing you to enjoy gaming fun anytime, anywhere.

&#x20;【Long-lasting Durability】With a built-in high-quality battery, it supports up to 30 hours of continuous gameplay, ensuring uninterrupted gaming sessions. Its highly wear-resistant surface treatment and robust structure ensure the controller can withstand daily wear and tear.&#x20;

【Professional-grade Performance】Whether for casual entertainment or competitive gaming, this controller delivers top-notch performance. Enhanced vibration feedback and programmable buttons add more fun and strategic depth to your gaming experience.

## **Preliminary Consideration**

**I.** Because the controller belongs to the gaming and technology product category, this can lead to concepts such as science fiction, high technology, and alien worlds. With these directions in mind, we can use AI to generate images and add inspirational materials. Open the SeaArt Canvas, select <mark style="background-color:yellow;">**Generate Canvas**</mark><mark style="background-color:yellow;">,</mark> enter the prompts in the image generation box, choose <mark style="background-color:yellow;">**4**</mark> for the number of images, and leave other options at default parameters. Click on Generate Images to obtain the generated pictures.

<figure><img src="/files/HPmRIUB7pgn3lw3yLE0z" alt="SeaArt Canvas - Controller"><figcaption></figcaption></figure>

**II.** In this mode, the generated images will remain on the canvas, allowing you to continue generating images by moving the image generation box aside, much like drawing on a sketchpad. You can delete unnecessary images and keep useful ones as references to more intuitively adjust image parameters.

<figure><img src="/files/uIjDncgtciyhEUI4j8Mb" alt="SeaArt Canvas - Adjust Image Parameters"><figcaption></figcaption></figure>

**III.** Once you have selected the reference images, we can proceed to the next step, which is to use the <mark style="background-color:yellow;">**Realtime Canvas**</mark> to add backgrounds to the product.

<figure><img src="/files/lZ7Wq6ias3GjDyTOGZ5f" alt="SeaArt Canvas - Realtime Canvas" width="563"><figcaption></figcaption></figure>

## **Realtime Canvas**

### **Step1: Upload the Product**

Click on the <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Product**</mark><mark style="background-color:yellow;">"</mark> option on the left side, and then select the <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Upload Product**</mark><mark style="background-color:yellow;">"</mark> function. After uploading the product image, the system will automatically perform the clipping operation to ensure that the product image remains clear and accurate during the drawing process.

<figure><img src="/files/POV9cmGQQtlR9zQORyUq" alt="SeaArt Canvas - Realtime Canvas - Upload the Product"><figcaption></figcaption></figure>

The scale adjustment tool is provided in the top left corner of the canvas, with a maximum output resolution of up to 1024x1024 pixels. Below that is the prompt box, where we can enter the prompts from the previous reference images.

<figure><img src="/files/TQPTISBOWoQN3wQCU5vp" alt="SeaArt Canvas - Realtime Canvas - Scale Adjustment" width="563"><figcaption></figcaption></figure>

### **Step2: Generate Background**

Next, let's start generating the background for the product. Select the <mark style="background-color:yellow;">**simple Pencil**</mark> tool from the toolbar, and use the dropdown options to adjust the brush size and color. The layer button at the bottom left allows users to adjust the order of each layer, flexibly managing the overlap between layers.

<figure><img src="/files/y5exh8HokGr9nCSJx0NK" alt="SeaArt Canvas - Realtime Canvas - Generate Background"><figcaption></figcaption></figure>

Before starting the generation process, adjust the relevant parameters such as <mark style="background-color:yellow;">**Denoising strength (recommended to set to 0.7), seed (-1)**</mark>, etc., to optimize the generation effect. Then, click on the <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Started/Stopped**</mark><mark style="background-color:yellow;">"</mark> button to begin the generation.

Adjust the final effect of the image by increasing or decreasing strokes and prompts. Once you have obtained the desired image, click on the download button in the top right corner to save the picture.

<figure><img src="/files/EaBHZ2GmsbFvLOENh0t5" alt="SeaArt Canvas - Realtime Canvas - Adjust Final Effect "><figcaption></figcaption></figure>

**Merge Background**

In addition to drawing the background manually, we can also directly use the previous reference images as the background. The SeaArt Canvas can seamlessly integrate the product with the background.

In <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Generate Canvas**</mark><mark style="background-color:yellow;">,"</mark> right-click on the selected image and choose <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Save to Mine.**</mark><mark style="background-color:yellow;">”</mark>

<figure><img src="/files/3wkMo78fCx8Ophz21e3u" alt="SeaArt Canvas - Realtime Canvas - Save to Mine"><figcaption></figcaption></figure>

Switch to <mark style="background-color:yellow;">**"Realtime Canvas,"**</mark> click on "Mine," select the saved image, and add it to the canvas.

At this point, you can <mark style="background-color:yellow;">**reduce the Denoising strength**</mark> to preserve more details of the original background image. As you can see, through the canvas, we can seamlessly integrate the product into the background, creating a natural and harmonious visual effect.

<figure><img src="/files/xI5IVh0XEcFimn3yXIqr" alt="SeaArt Canvas - Realtime Canvas - Reduce the Denoising Strength"><figcaption></figcaption></figure>

### **Step3: Enhancing Image Quality**

Due to the size limitation of the canvas, we also need to enhance the image quality of the final picture.

Open the <mark style="background-color:yellow;">**Swift AI**</mark> and select the <mark style="background-color:yellow;">**AI Image Upscaler.**</mark>

Import the image that needs enhancement and adjust the specific parameters according to your requirements.

<figure><img src="/files/XL6GAAKXywKmQfLN0Zcx" alt="SeaArt Canvas - Realtime Canvas - Adjust the Specific Parameters"><figcaption></figcaption></figure>

### **Step4: Final Result**

Finally, edit the enhanced image and add necessary text descriptions. Our e-commerce poster is now complete.

<figure><img src="/files/nGfFsVzXiTjz6FhPqdM3" alt="SeaArt Canvas - Final Result"><figcaption></figcaption></figure>

With the continuous advancement and innovation of AI technology, future versions of SeaArt Canvas will introduce more intelligent features, such as deeper image analysis and generation algorithms. This will further streamline the design process and enhance creative efficiency.


# Composite Poster

Here, we will guide you step-by-step on how to create a stunning composite poster using the SeaArt Canvas.

Whether you're just starting out with AI art or you're a seasoned designer, you're bound to discover new techniques and find fresh inspiration in this process.

<figure><img src="/files/zaTWdJzFTLBg16YoY3b2" alt=""><figcaption></figcaption></figure>

## STEP 1 Open SeaArt Canvas

1. Open SeaArt Canvas and select the <mark style="background-color:yellow;">'Realtime Canvas.'</mark> SeaArt Canvas offers 'Generate Canvas,' 'Realtime Canvas,' and 'Design Canvas.'"

* Realtime Canvas allows you to view and adjust images in real-time.
* Generate Canvas enables direct use of AI to generate images, perform partial redraws, and erasing, among other functions.

<figure><img src="/files/YX2hDacGyiKSKVssYnzd" alt=""><figcaption></figcaption></figure>

## STEP 2 Determine the Poster Theme

Before we begin, we need to clarify the theme of the poster we want to create. For instance, our theme could be <mark style="background-color:yellow;">"Being Chased by a Giant Cat."</mark> With a theme in mind, we can use the generation mode to produce several sketches as references, providing us with a general creative direction.

* **Prompt:** A huge ferocious long-haired cat running, chasing a man, disaster film.

<figure><img src="/files/bXUI7dtcmDPihNmCI1NJ" alt=""><figcaption></figcaption></figure>

As you can see, the directly generated image does not meet our needs, and there is a significant gap between the actual result and our envisioned effect.

## STEP 3 Generate Poster Materials

Next, generate the materials needed for the poster individually and save them as cutouts.&#x20;

Open the Realtime Canvas, and first adjust the image generation parameters, selecting the appropriate model and lora.

* Model: RealVisXL V4.0
* Lora: Movie Aesthetic

<figure><img src="/files/Ow25q3xf9DWJ96yIDxmi" alt=""><figcaption></figcaption></figure>

The two sampler options in the settings control the generation mode and Realtime Canvas respectively. We will change the sampling method in Realtime Canvas **from LCM to DPM++ 2M Karras.** Although this will increase the image generation time, it enhances the quality of the generated images.

<figure><img src="/files/K1bD3U7ALruiDawxsmQE" alt=""><figcaption></figcaption></figure>

Once the parameters are set, we directly enter text into the prompt box to generate the cat material images.

* **Prompt:** A ferocious long-haired cat running, cute, huge cat, low pov.

<figure><img src="/files/mAGkRKOBPaEDjwRppZRJ" alt="" width="428"><figcaption></figcaption></figure>

After selecting the appropriate material image, click <mark style="background-color:yellow;">**"Save as New Layer"**</mark> in the top right corner to copy the image to the generation area. Remember to click the **pause button** during this operation.

<figure><img src="/files/fbh4Qeqto78pB1ewA0b6" alt="" width="434"><figcaption></figcaption></figure>

Next, click on the image and select the <mark style="background-color:yellow;">**"Remove Background"**</mark> feature to cut out the cat.

<figure><img src="/files/zB9SWpWDtMMxnIC9DVjk" alt="" width="464"><figcaption></figcaption></figure>

Then, click the layers button in the bottom left corner to open the layers panel, and **click the little eye icon to hide the cat layer.**

<figure><img src="/files/00WAgO2GxwjlseTk1ShH" alt="" width="563"><figcaption></figcaption></figure>

Repeat the steps above to continue saving other materials to the generation area.

<figure><img src="/files/y3yvY7UAL8UzzPfuaC68" alt="" width="292"><figcaption></figcaption></figure>

> **Prompt:** A terrified man, running in fear, low POV, white background

<figure><img src="/files/tOAWJR2NaWQNzLB8q4ef" alt="" width="303"><figcaption></figcaption></figure>

> **Prompt:** Darkness, leaf-covered clearing, jungle, path, huge dust and smoke, low perspective

## STEP 4 Generate Images

Adjust the size and position of the materials to align with the perspective.&#x20;

<figure><img src="/files/HZdgkevztmzy06lfEDKy" alt="" width="563"><figcaption></figcaption></figure>

Enter the scene prompt, describing the content of the image as completely as possible.

<figure><img src="/files/cdDV0NvkdP5nJHdWJL6T" alt=""><figcaption></figcaption></figure>

> **Prompt:** Cute giant cat running fast, chasing a man who is running, scared man, speed blur, smoke filled, rock splash, wide angle lens, motion blur, film lighting, Sony camera, backlighting, edge lighting, ray tracing, film grain

After generating a satisfactory image, click the <mark style="background-color:yellow;">**download**</mark> button in the top right corner to save the image. Finally, it's recommended to use <mark style="background-color:yellow;">**AI Image Upscaler**</mark> to add detail to the scene.

<figure><img src="/files/4YQCpzyZUjKk2NDAq7rY" alt=""><figcaption></figcaption></figure>

The final product effect is as follows:

<figure><img src="/files/WRnGdjSiIs0b6rT1vEDy" alt="" width="375"><figcaption></figcaption></figure>

Using the SeaArt Canvas allows you to more clearly adjust the image to generate the perfect picture in your mind!


# 4-Parameters

This quick parameters guide shows you how to master AI image and art generation.


# 4-1 Model

Unlock AI image creation with SeaArt! Explore Checkpoints, Loras, and their interplay to generate stunning and unique visuals.

> **In SeaArt, there's no longer a need for local model deployment. You can simply choose online, or even upload and train your own models.**

**Creation Process:** <mark style="background-color:red;">Choose model</mark> - Enter prompts - Adjust relevant parameters - Generate

## Checkpoint

Generally speaking, both Lora and Checkpoint are referred to as models. Therefore, we can call Checkpoint the large model or base model. Compared to Lora, the large model is relatively larger, ranging from 2GB to 7GB, and it determines the <mark style="background-color:yellow;">main style</mark> of the final image.

Same prompts, different Model:

<mark style="background-color:yellow;">Prompts: black hair, princess, frills, standing, short hair</mark>

**Recommended Checkpoint:**

Realistic:

Anime:

## LoRA

LoRA is typically around 100MB in size. It can "fine-tune" the style of images, fix the appearance, posture, etc. On top of overlaying the Checkpoint, adjusting weights can achieve different effects. In the SeaArt, Lora can stack up to 5.

**Character LoRA:** These models specialize in capturing specific character traits, such as appearance, body proportions, and expressions, commonly found in cartoons, video games, or other media. They prove valuable for fan artwork creation, game development, and animation/illustration projects.

**Style LoRA:** These models are tailored to mimic the artistic style of specific artists or artistic movements. They excel in transforming reference images into the desired aesthetic, making them useful for stylization purposes.

**Clothing LoRA:** Similar to Style LoRA, Clothing LoRA models focus on specific clothing styles or fashion aesthetics. They are adept at generating images with a particular clothing aesthetic, offering applications in fashion design and digital styling.

Different Lora with the same Checkpoint and prompt:

Prompt: A girl, frontal view, close-up, sitting on a plush chair, exquisite face, full of confidence, luxurious atmosphere, top quality, 8k.

Add multiple Loras:

After adding the Team Rocket Uniform, not only did it increase the weight of Lora, but it also added related prompts for Lora: <mark style="background-color:yellow;">"Team Rocket uniform," "red letter R," "white skirt," "white short top," "black long sleeves," "black elbow gloves,"</mark> making the effect of Lora more pronounced.

> \
> [Previous4-Parameters](https://app.gitbook.com/o/SXzpAogddQhmphNiXJbH/s/igAtVLBrlaI8jVruJfC8/~/edit/~/changes/212/4-parameters)[Next4-2 Mode](https://app.gitbook.com/o/SXzpAogddQhmphNiXJbH/s/igAtVLBrlaI8jVruJfC8/~/edit/~/changes/212/4-parameters/4-2-mode)SULast modified 1yr agoOn this page[Checkpoint](https://app.gitbook.com/o/SXzpAogddQhmphNiXJbH/s/igAtVLBrlaI8jVruJfC8/~/edit/~/changes/212/4-parameters/4-1-model#checkpoint)[LoR](https://app.gitbook.com/o/SXzpAogddQhmphNiXJbH/s/igAtVLBrlaI8jVruJfC8/~/edit/~/changes/212/4-parameters/4-1-model#lora)If the effect is not significant after adding Lora, we can try:&#x20;
>
> I. Adjusting the weight of Lora&#x20;
>
> II. Entering corresponding prompts

Compared to Checkpoint, Lora has shorter training time and higher "flexibility," and can exert good effects on controlling images.

<mark style="color:red;">Attention!!!</mark>

I. SD1.5 and SDXL's Lora cannot be used together.

II. Most Lora require trigger words. Including these words in the prompts can emphasize the unique subject or style provided by the Lora. These trigger words are not always necessary, but it is recommended to include them in the prompts to ensure stable image generation.

**Different Lora Effects:**

1. **Detail Tweaker LoRA:** The higher the weight, the more details.

<figure><img src="/files/2n3SGgHcf2SMzU15bNmU" alt="AI Image Generation Model - Detail Tweaker LoRA" width="563"><figcaption></figcaption></figure>

2. **epi\_noiseoffset**: Increases contrast.
3. **Makima (Chainsaw Man) LoRA**: Fixes the character's image.
4. **ankymoore illustration**: Modifies the art style.

<mark style="color:red;">\*</mark>*Lora has many functions, trying different Loras may bring more surprises!*


# 4-2 Mode

SeaArt simplifies AI art! Choose from default, automatic, or advanced modes to easily generate stunning images with powerful models.

**Creation Process:** Select Model - Enter Prompts - <mark style="background-color:yellow;">Adjust Parameters</mark> - Generate

**Default**: The most common image generation mode, which can be combined with HD restoration.

**SeaArt 2.0/2.1** has a better understanding of prompts, resulting in higher image quality. Users can choose different modes according to their needs.


# 4-3 Basic Settings

Control your AI art in SeaArt! Adjust image quantity, mode, and size to achieve your desired level of detail and composition.

**Creation Process:** Select Model - Enter Prompts - <mark style="background-color:yellow;">Adjust Parameters</mark> - Generate

Set the **Quantity, Mode, and Size** of images:

**Image Quantity**: How many images to generate at once.

**Image Mode**: The clarity of the images.

**Image Size**: Typically, portrait orientation is used for generating character images, while landscape orientation is used for scenery images. Maximum supported size is 1024\*1204 pixels.

> I. When the resolution is constant, larger image sizes can accommodate more information, resulting in richer details in the final image. If the image size is set to be small, the image content may appear rough.&#x20;
>
> II. Although selecting a larger image size can enhance the richness of details, it also increases computational burden, prolongs the time required for drawing, and raises computational resource consumption. Additionally, when the image size exceeds a certain range, the model interprets it as multiple images stitched together, resulting in issues such as multiple people or limbs appearing within one image.


# 4-4 Advanced Config

Fine-tune your AI art in SeaArt! Master negative prompts, VAEs, sampling, CFG scale, seed, and clip skip for precise control.

**Creation Process:** Select Model - Enter Prompts - <mark style="background-color:yellow;">Adjust Parameters</mark> - Generate

### Negative Prompts

Generally, it's challenging for the model to understand negations in prompts, such as words like "no," "not," "except," or "without." Therefore, we need to include unwanted effects in the Negative Prompts. Apart from adding elements we don't want in the image, we can also include words like "low quality," "low detail," "ugly," "deformed," etc. This helps improve the quality of the final image. Typically, when generating images, SeaArt automatically includes negative labels.

> (worst quality, low quality, normal quality, lowres, low details, oversaturated, undersaturated, overexposed, underexposed, grayscale, bw, bad photo, bad photography, bad art:1.4), (watermark, signature, tet font, username, error, logo, words, letters, digits, autograph, trademark, name:1.2), (blur, blurry, grainy), morbid, ugly, asymmetrical, mutated malformed, mutilated, poorly lit, bad shadow, draft, cropped, out of frame, cut off, censored, jpeg artifacts, out of focus, glitch, duplicate, (bad hands, bad anatomy, bad body, bad face, bad teeth, bad arms, bad legs, deformities:1.3)

### VAE

VAE can be regarded as a kind of "filter" that improves the quality of image generation and enhances visual effects through optimization algorithms. It can also make slight adjustments to the shapes of images. If you notice color issues in the images, you can try switching to a different VAE.

**Commonly used VAEs:**

**Automatic:** Automatically selects the most suitable VAE configuration for the current task.

**None:** Does not use any VAE.

**vae-ft-mse-840000-ema-pruned:** Realistic color style, 840000 indicates the number of training iterations, which helps improve the quality of generated images while reducing complexity and increasing efficiency.

**vae-ft-ema-560000-ema-pruned:** Realistic color style, trained for 560000 iterations, suitable for faster or lower resource-consuming image generation.

**kl-f8-anime2:** Optimized for generating images in anime style.

<mark style="color:red;">\*</mark>Some Checkpoint come with built-in VAE, so there is no need to select VAE separately.

### Sampling

#### **Sampling principle**

A standard AI painting process typically involves forward addition of noise and backward denoising, restoration, and target generation. During the forward process, noise is continually added to the input data, while the sampler is responsible for denoising during the backward process.

<figure><img src="/files/h13x3x0OgClKo9gTudf8" alt="SeaArt AI - Advanced Config - Sampling - Sampling Principle" width="563"><figcaption></figcaption></figure>

Forward process (from right to left): Gradually adding noise to the original image, mainly during the training process to train the U-Net network's ability to predict noise.

Backward process (from left to right): Gradually denoising the estimated noise by the trained U-Net network, ultimately reproducing the image.

In these two processes, AI effectively scrambles a specific image, then learns from parts of it to create a new image in reverse. That is, once the forward process is trained, the backward process generates a completely new image from a noisy image.

Before generating a clear image, the model needs to generate a random image in the latent space. The noise predictor starts working by subtracting the predicted noise from the image. With repeated steps, we eventually obtain a clear image. The entire denoising process can be referred to as "sampling," and the method used in sampling is called a sampler or sampling method.

The sampling method determines how the denoising is performed, and different sampling methods yield different image results.

\*The numerical value of the seed determines the initial noise of the first generated image.

<figure><img src="/files/DndAj5461chAItQxJH6z" alt="Seed Determines the Initial Noise of the First Generated Image" width="240"><figcaption></figcaption></figure>

#### **Sampling Method**

1. **Old-School ODE Solvers**

Euler: Euler method, the simplest solver.

Heun: More accurate but slower version of the Euler method.

LMS: Linear Multistep Method, same speed as Euler but more accurate.

<mark style="background-color:yellow;">Convergence:</mark> As the number of sampling steps increases, the sampled results eventually tend toward a <mark style="color:red;">fixed image</mark>, and the image gradually stabilizes.

2. **Ancestral Samplers (names include an "a")**

Euler a

DPM2 a

DPM++ 2S a

DPM++ 2S a Karras

These samplers add noise at each sampling step, thus exhibiting a degree of randomness and not converging.

<mark style="background-color:yellow;">Non-convergence:</mark> The images are <mark style="color:red;">random</mark> and may add some details. To obtain stable and reproducible results, one should avoid using ancestral samplers.

<mark style="color:red;">\*</mark>Some samplers, even without an "a" in their name, are also random samplers.&#x20;

3. **DDIM, PLMS (no longer widely used)**

DDIM: Denoising Diffusion Implicit Models, the first sampler designed for diffusion models.

PLMS: Pseudo Linear Multistep Method, a faster alternative to DDIM.

4. **DPM and DPM++ Series**

These samplers have a high utilization of tags, appropriately magnifying sampling steps for better effects, though overall speed is slower. DPM++ is an improvement over DPM, yielding more accurate results but at a slower pace.

Karras: Produces clear images with fewer sampling steps, optimizing the algorithm.

5. **Restart:** Uses fewer sampling steps to generate good images in less time.

LCM: Generates images quickly.

<mark style="color:red;">\*</mark>**Recommended to use:**

* Euler/Euler a: Fast speed, high quality, suitable for most scenarios, recommended steps are 15-30.
* <mark style="color:red;">DPM++2M Karras: Convergent, fast speed, good quality (15-25 steps).</mark>
* DPM++SDE Karras: Not convergent, slow speed, good quality, suitable for realistic images, recommended 10-15 steps.
* DPM++2M SDE Karras: Intermediate algorithm between 2M and SDE, not convergent, slightly faster speed.
* DPM++ 2M SDE Heun Exponential: Not convergent, soft and clean image, with fewer details.
* DPM++ 3M SDE Karras
* DPM++ 3M SDE Exponential: Same speed as 2M, requires more sampling steps, when sampling steps > 30, lower the text intensity (CFG) for better results.
* Restart: Very fast speed, only suitable for quickly producing drafts or concept verification, ideal results can be achieved with very few steps.
* LCM: "Real-time rendering" can be achieved in only 4 steps, although the image quality is average, it is suitable for generating inspirational sketches or preliminary concept designs.

<mark style="color:red;">\*Note</mark>

<mark style="color:red;">I. Prioritize the algorithms recommended by the model author to ensure the best compatibility and effectiveness.</mark>

<mark style="color:red;">II. Prioritize using algorithms with a plus sign, as optimized algorithms tend to be more stable than those without a plus sign.</mark>

<mark style="color:red;">III. When encountering noise issues in the generated images, consider trying a different sampler.</mark>

#### **Sampling Steps**

Generally, the higher the number of sampling steps, the better the quality of the image. However, around 25 sampling steps are usually sufficient to achieve high-quality images. Increasing the number of steps beyond this point may generate different images, but it doesn't necessarily guarantee better quality. Additionally, higher sampling steps require more time. In most cases, there's no need to set excessively high sampling steps, which would only increase the waiting time.

As the number of sampling steps increases, the main form of the "girl" remains relatively consistent, while certain small details such as hair quality, color, background, etc., improve with the increase in steps. Therefore, the number of sampling steps should be adjusted according to one's own needs.

### CFG Scale

The relevance to the prompts: The higher the text strength, the closer the image is to the prompts. It's generally set around 7-10. If it's set too high, it may cause image breakdown. If the generated image does not follow the prompts, you can increase the text strength appropriately.

> Prompts: full lenght shot, super hero pose, biomechanical suit, inflateble shapes, wearing epic bionic cyborg implants, masterpiece, intricate, biopunk futuristic wardrobe, highly detailed, artstation, concept art, cyberpunk, octane render

### Seed

During the drawing process, there is significant uncertainty in AI Art , as each drawing involves a set of random computational mechanisms, each corresponding to a fixed seed value. By fixing the seed value, we can control the randomness of the drawing results.

For example, if we are satisfied with a particular image generated, we can fill in its seed value here to reproduce the same content. Clicking 'Random' resets the seed to the default -1, while 'Customization' allows you to freely enter the seed value.

Using the same parameters, prompts, and seed will produce identical images. Therefore, we can utilize the same seed to modify certain parameters, resulting in the generation of new images with the original features.

<mark style="color:red;">\*</mark>Only modifying the emotional words to change the facial expression while keeping other features such as <mark style="background-color:yellow;">hair, clothes, and background</mark> unchanged.

### Clip Skip

Layer by layer, the prompts are transformed into numbers, then read by the converter, providing a progressively detailed understanding of the prompts.

If the prompt is: "A young girl, wearing a black dress, with a black hat, holding a wand, a witch," when Clip Skip is set to 2, the AI may omit the concepts of the black dress or the wand. <mark style="background-color:yellow;">As the Clip Skip value increases, the AI will omit more of the prompt.</mark>

Therefore, when Clip Skip is set to 1, it means terminating the image from the last layer. The result will include a complete description of the prompts. The earlier the termination, the less description will be obtained from the prompts, resulting in less accuracy in the final result. It's generally set to 2.

**What is Clip Skip for?**

Clip Skip helps to <mark style="background-color:yellow;">address overfitting</mark> situations by terminating the reading of the prompts in a timely manner. When the image is overfitted, Clip Skip can be increased.

By setting Clip Skip, you can <mark style="background-color:yellow;">adjust the details and style</mark> of AI Art, making the final result more flexible and controllable, thus meeting different image generation requirements.

> Prompts: best quality,masterpiece,illustration,beautiful detailed glow,textile shading,absurdres,highres,dynamic lighting,intricate detailed,beautiful eyes,\[backlighting],face lighting,(pov:1.3), (1 girl, solo:1.5),asymmetric bang,black hair,(smile),(jeans pants and shirts)


# 4-5 Advanced Repair

Enhance your AI art with SeaArt's advanced tools! Upscale for higher resolution, and repair characters and faces for flawless results.

### Upscale

Enabling it can increase the output image size and quality, but it will also increase the rendering time and computational resources consumption. The resulting image with enhancements will have richer details and higher visual quality.

Algorithm and parameters:

**4x-UltraSharp:** This algorithm increases the image resolution by 4 times. It pays special attention to <mark style="background-color:yellow;">maintaining the sharpness of edges and details when enlarging images.</mark>

**R-ESRGAN 4x+:** This algorithm increases the image resolution by more than 4 times. It uses GAN to improve image quality, resulting in <mark style="background-color:yellow;">more realistic images with finer textures.</mark>

**R-ESRGAN 4x+ Anime6B:** This algorithm increases the image resolution by more than 4 times and is optimized for <mark style="background-color:yellow;">anime-style images.</mark>

**8x\_NMKD-Superscale\_150000\_G:** This algorithm increases the image resolution by 8 times and is suitable for <mark style="background-color:yellow;">applications requiring high-resolution output.</mark>

Other options are generally set to their default values.

### Character Repair

Mainly used for fixing issues with characters' hands, faces, and bodies, essentially incorporating an automatic partial Repainting function.

Prompts:

Face: Exquisite face…

Hands: Normal hands, five fingers…

Denoising Strength: Similar to the extent of redrawing; the larger the value, the greater the changes after generating the image.

Confidence: The range of recognition for the parts; the smaller the value, the smaller the recognition range. It is recommended to lower the value for scenes with multiple people

### Restore Faces

Use this option to repair facial details and prevent the image from collapsing too much. When the face occupies a large proportion of the image, selecting this option may result in excessive fitting and blurring. It is recommended to select this option when the face is in a long shot. Use this with real photos.


# 4-6 Complete Prompting Guide

This guide will help you to understand every details about prompts.

> Prompting is first and most important phase of AI image generation. Good prompting skills will help users to acquire great generations on every model.&#x20;

## Building With Categories And Order

Good prompt requires correct order and categorization. It could look same but **prompt placements actually matters**. General Stable Diffusion order and categories looks like this:

1. View/Shot
2. Subject (with details)
3. Medium
4. Resolution/Detail Terms
5. Style References (includes artist references)
6. Additional details (background, environment etc.)
7. Main Color Themes
8. Lighting/Shadow
9. Quality Terms `trending on deviantart, (best image:1.5)` etc.

Howover, order and categorization can show difference for each model. You can see them on model guides. But this can help you on every model.

You don't have to use every category in your prompt.

Categorize every term in your prompt. For example, "subject, facial details, body details, dress details" etc.

**Wrong:** 1girl, ginger, blue eyes, smiling, red dress, wavy hair, thin&#x20;

**Correct:** 1girl, smiling, blue eyes, ginger, wavy hair, thin, red dress This will reduce wrong generations most of time.

## Negative Prompts

Negative prompts will help you avoid unwanted stuff in your generations. Main subjects for negative prompts:

* Negative Quality Prompts: worst quality, worst image, low quality, lowres etc.
* Embeddings: verybadimagenegative\_v1.3, ng\_deepnegative\_v1\_75t, badhandv4 etc.
* Unwanted stuff: NSFW, flower, rain, hat etc.

## Prompt Weighting

Weights are an Stable Diffusion feature to control generated image. `(term:factor)` , `term` is what you want to increase or decrease in image, `factor` is scale to control it.&#x20;

**term:** Term can be everything you want to edit in your prompt.&#x20;

**factor:** Factor actually can scale between `0.01-100` but this type weights create deformity in your images. Universal recommended scale for weighting is `0.5-1.5`. **Keep your weights between this scale.** Howover, some models can handle up to `2` weight.

### What Does Weight, Scale Mean?

Weights are how much attention you put on AI for that spesific term at each step. They are absolute, this means increasing value of one term will not reduce weights of other terms.

**Note:** `1girl, brunette, garden` is not same with `(1girl:1.5), (brunette:1.5), (garden:1.5)`, they have different work weight and outputs will be different than each other.

Scale is based on percentage system, which is mean `1 = %100, 1.5 = %150`. This is why you can't go with values higher than `1.5`. AI can't handle this type work weight for one term, because of their training.

### Different Types of Weighting

(term:factor)

`factor` can scale between `0.5-1.5` and this type can be used for both reducing and increasing weights. `[]` is not working with this system. Use only `()`, if you are going to specify value.

`1girl, (red hair:1.2), (rose pupils:0.7)`

#### (term)

* Each `()` increases value of term by x1.1

(term) = (term:1.1) ((term)) = (term:1.21) (((term))) = (term:1.331) ...

#### \[term]

* Each `[]` reduces value of term by x1.1

\[term] = (term:0.9) \[\[term]] = (term:0.82) \[\[\[term]]] = (term:0.75) ...&#x20;

**Note:** The maximum recommended number of brackets is 3, values higher than 3 can cause deformity in images.

## Step-by-Step Prompting

I will show an example of writing prompt step by step here, will use DreamShaper v8 Model because it can handle various prompt orders (including general one) and it's beginner friendly model. All images were created with same settings.&#x20;

Sampler: DPM++ 2M KARRAS&#x20;

CFG: 9&#x20;

Sampling Steps: 30

### Choosing Subject And Main Theme

Let's say, we want to create a illustration of a girl in space. Starting with `illustration, 1girl, space` and now we have our first output.

<figure><img src="/files/ewWDWdUR8bFiyUM9lDWS" alt="" width="384"><figcaption></figcaption></figure>

### Negative And Embeddings

Since we have a subject, now we can write negative prompts and embeddings to maintain quality on each step. I will go with a few basic negative and DreamShaper Embedding.

**Positive Prompt**

> illustration, 1girl, space

**Negative Prompt**

> BadDream, (UnrealisticDream:1.3), low res, NSFW, grayscale, monochrome, astronaut suit, nude, child, kid

### Adding Medium And View

After reaching acceptable quality, time to add medium and view terms to get consistent results. I will go with `digital art` and `watercolor ink` also add `mixed media` to get mixture of these mediums. As a view, `upper-body illustration` is good choice for us.&#x20;

**Positive Prompt**

> upper-body illustration, 1girl, (digital art), watercolor ink, space, mixed media

### Improving quality

Since we have good base, we can improve quality with a few word.&#x20;

**Positive Prompt**

> upper-body illustration, 1girl, (digital art), watercolor ink, (ultra quality:1.3), masterpiece, (highly detailed), HDR, 8K, space, mixed media

### Adding Style With References

I added two artist reference and abstract art with sci-fi as a theme, because i love portraits of Charlie Bowater and Maciej Kuciara has good sci-fi character desings.&#x20;

**Positive Prompt**

> Positive Prompt upper-body illustration, 1girl, (digital art), watercolor ink, (ultra quality:1.3), masterpiece, (highly detailed), HDR, 8K, abstract art, (by Charlie Bowater:0.9), (style of Maciej Kuciara), sci-fi, space, mixed media

### Additional Details And Colors

We can add details we want and choose color themes before final stage. I added details to my subject and background also chose red and black as a main colors. I used `helmet` on negative because i want to see her face.

**Positive Prompt**

> upper-body illustration, 1girl, black hair, bobcut, futuristic space suit, (digital art), watercolor ink, (ultra quality:1.3), masterpiece, (highly detailed), HDR, 8K, abstract art, (by Charlie Bowater:0.9), (style of Maciej Kuciara), sci-fi, fractal stars, space, amazing space background, red theme, black theme, mixed media

**Negative Prompt**

> BadDream, (UnrealisticDream:1.3), low res, NSFW, grayscale, monochrome, astronaut suit, nude, child, kid, helmet

### Lighting And Final Touches

Now we can finish our prompt with focus, light, shadow and last quality touches. I also added a few negative for better faces.

**Positive Prompt**

> upper-body illustration, portrait, 1girl, black hair, bobcut, futuristic space suit, (digital art), watercolor ink, (ultra quality:1.3), masterpiece, (highly detailed), HDR, 8K, abstract art, (by Charlie Bowater:0.9), (style of Maciej Kuciara), sci-fi, fractal stars, space, amazing space background, red theme, black theme, (depth of field), sharp focus, diffused lighting, backlighting, soft shadows, mixed media, trending on deviantart, award winning art

**Negative Prompt**

> BadDream, (UnrealisticDream:1.3), (deformed eyes:1.3), ugly, low res, NSFW, grayscale, monochrome, astronaut suit, nude, child, kid, helmet

**Now you learned how to create an effective prompt can generate great results most of time. Follow this guidelines, create your art and don't forget to share with**[ **SeaArt** ](https://www.seaart.ai/)**Community.**


# 4-7 Prompt Edit | Keyword Blending Guide

This guide explains advanced prompting techniques for better control over image generation on SeaArt.AI, applicable to all models except Flux, which uses a different system.

You have learned how to write basic prompt with weights on ⁠Complete Prompting Guide. This guide will teach you advanced prompting techniques to have more control over image generation. Effects can show difference for each model but these rules are applies to all models work on [SeaArt.AI](https://docs.seaart.ai/guide-1/4-parameters/www.seaart.ai). Only Flux doesn't work with these tools because it has different prompt understanding system.

## Before Start

* Term `keyword` can be any word in your prompt. Keyword can be more than one word in some syntax.
* Term `factor` is value of sampling steps. Scales between 1 and highest sampling step.
* Term `factor%` is percentage value of sampling steps. Can't be higher that 1 (100%) and scales between 0-1.

## \[keyword1:keyword2:factor] Syntax

Generation starts with `keyword1` and changes to `keyword2` after `factor` step.

* Lower values sometimes will not have visible effect because it has small part of generation.
* Order matters, first `keyword1` is your starting term while `keyword2` is secondary part of generation.

### Example

`portrait of a girl, [smiling:frown:0.7] face` with 20 step generation will be generated as:&#x20;

**Step 1-14** `portrait of a girl, smiling face`&#x20;

**Step 15-20** `portrait of a girl, frown face`

### Below You Can See Difference

Images were generated with same settings with 20 steps. `factor%` changed with `0.1, 0.3, 0.5, 0.7, 0.9`&#x20;

As you can see first image has more angry face because only 2 step generated with `smiling` while last image has smiling face because 18 step generated with `smiling` and changed into `frown` at last 2 steps.

<figure><img src="/files/AMjL0vaDw336pIHunvkk" alt=""><figcaption></figcaption></figure>

## \[keyword:factor] Syntax

Generation starts without `keyword` and adds `keyword` to prompt after `factor` step.

* Higher values will not have visible effect because main composition of image will be generated already and AI cannot make this type significant changes.

### Example

`portrait of a girl, green eyes, [pink hair:7]` with 20 step generation will be generated as:&#x20;

**Step 1-7** `portrait of a girl, green eyes,`&#x20;

**Step 8-20** `portrait of a girl, green eyes, pink hair`

### Below You Can See Difference

Images were generated with same settings with 20 steps. Term `factor` is changed with `3, 7, 10, 13, 17`&#x20;

As you can see, there is no pinkish hair in the third and later images because AI already generated main composition and couldn't make changes even we add `pink hair` later. White color was totally luck, it could be any color includes pink.

## \[keyword::factor] Syntax

Generation starts with `keyword` and removes `keyword` from prompt after `factor` step.

* Higher values will not have significant differences because main composition will be generated with `keyword` already.

### Example

`portrait of a girl, [ginger::3], black dress` with 20 step generation will be generated as:&#x20;

**Step 1-3** `portrait of a girl, ginger, black dress`&#x20;

**Step 4-20** `portrait of a girl, , black dress`

### Below You Can See Difference

Images were generated with same settings with 20 steps. Term `factor` is changed with `3, 5, 7, 10, 13`&#x20;

As you can see there is no significant difference after third image because main composition was already generated when we remove `keyword` from prompt. Blonde hair was totally luck, it could be any hair color includes ginger.

## \[keyword1|keyword2] Syntax

This syntax will swap `keywords` on each step. `keyword` can be more than two, you can add `keyword` as much as sampling steps but it's not recommended.

* Lower `keyword` count equals to better results.

### Example

`portrait of a girl, [blonde|ginger|black] hair` with 20 step generation will be generated as:&#x20;

**Step 1-4-7-10-13-16-19** `portrait of a girl, [blonde|ginger|black] hair`&#x20;

**Step 2-5-8-11-14-17-20** `portrait of a girl, [blonde|ginger|black] hair`&#x20;

**Step 3-6-9-12-15-18** `portrait of a girl, [blonde|ginger|black] hair`

### Below You Can See Examples

1. `portrait of a girl, [blonde|ginger|black] hair`
2. `[frog|bird|cat]`

<figure><img src="/files/LgVD0ucXqrQeBdaPo4tr" alt="" width="563"><figcaption></figcaption></figure>

## keyword1 (keyword2) Syntax

Slash can be used to use actual parantheses in prompt without weight. Some tags include parantheses `()` and this syntax help you to use them.

### Example

`mona (genshin impact)` will be tokenized as a `mona` with 1x weight and `genshin impact` with 1.1x weight.&#x20;

`mona \(genshin impact\)` will be tokenized as itself, without weighting and with actual parantheses.

### Below You Can See Example

`(masterpiece, best quality), mona \(genshin impact\)`


# 5-Practical Examples

Boost your creativity and simplify your life with SeaArt AI! Explore practical examples and discover how SeaArt AI can empower you.


# AI Influencer

Create AI influencers and monetize them! Learn how to design, manage, and profit from virtual influencers on social media.

**How to Create AI Influencers with** [**SeaArt.AI**](http://SeaArt.AI) **and Generate Profit?**

From brand endorsements to magazine covers, an increasing number of Virtual Influencers are appearing on social media and in the fashion industry, attracting many fans and making profits for their creators.

## **1.** **Successful Cases**

### **Lil Miquela**

The top North American virtual influencer, Lil Miquela, with 2.6 million followers, enjoys sharing her fashion, food, and romantic life on Instagram. Currently, Lil Miquela charges **$8,500** for endorsements and has successfully collaborated with famous brands like CHANEL, Prada, Vetements, Supreme, etc. According to data from the UK company OnBuy, in 2020 alone, Lil Miquela generated revenue of $11.7 million for her managing company Brud.

<figure><img src="/files/xFuE9uUnyfulUPsxR0W8" alt="Lil Miquela on Instagram"><figcaption></figcaption></figure>

### **Emily Pellegrini**

Emily Pellegrini's rapid rise to fame on Instagram has caught the attention of soccer stars, billionaires, and celebrity athletes, now boasting almost 300,000 followers. Within six weeks of her debut, Emily Pellegrini has earned **nearly $10,000** in the private domain and is touted as the world's most popular artificial intelligence.

### <mark style="background-color:yellow;">**SeaArt Monetization Cases**</mark>

**Ambra Fontana**

**Ambra Fontana**, a virtual influencer created with [SeaArt.AI](http://SeaArt.AI) by our user, has gained nearly 9,000 followers on Instagram in just one month. The account's combined value reaches **$1,000** **per month**, with sales of exclusive content on private domain platforms alone surpassing monthly earnings of **$500**. The creator told us the workflow **relies entirely on** [**SeaArt.AI**](http://SeaArt.AI) **generation** without the need for additional tools.

💡 <https://www.instagram.com/ambrissima/>

**Alert: On instagram there are many pages that offer to increase followers by paying between 30$ and 200$. DON'T TRUST. the followers who fall back on you are all bots, this will only lead to your profile having fake followers as well as spending money unnecessarily.**

A useful method to increase followers is to ask for collaborations from much larger pages. the maximum number of collaborations you can ask for a single post is 5. sometimes these pages ask you to pay a fee, see if it makes sense. a useful method to understand if a page is trustworthy is to compare followers - likes of the posts. example: if a page has 100k followers but its posts only have 100-300 likes, it means that they are all bot and it doesn't make sense to pay. if instead the posts have more than 2k likes, it may make sense to pay a fee. keep in mind, however, that there are always many pages that do free collaborations, just look for them! (From creator)

## 2. **Monetization Strategies**

AI Virtual influencers can be monetized in the following ways:

### **Brand Promotion**

This involves <mark style="background-color:yellow;">accepting advertisements on social media platforms.</mark> Currently, numerous AI influencers are making profits this way, especially on platforms like Instagram and Twitter.

### **Brand Endorsements**

After gaining more fans and traffic, some brands will take notice and invite virtual influencers to collaborate and <mark style="background-color:yellow;">endorse their products</mark>, generating income in the process.

### **Private Domain Sales**

This is by far the most common monetization method for virtual influencers. The creator can place a "one link" on their social platform and create massive exclusive content (such as **customized photos** and **paid conversations**) within their private domain. This content is <mark style="background-color:yellow;">**only visible**</mark> after payment, thus drawing in fans willing to spend.

## 3. **Creation & Management of AI Virtual Influencers**

### **Define Persona**

The character can be a girl, boy, infant, pet, etc. Add a name, personality, and identity (such as student, idol, housewife, CEO, supermodel, and fashion blogger). Choose the persona you think is most suitable or capable of monetization through this method.

### **Create Image**

An important aspect of character creation is to have a <mark style="background-color:yellow;">**consistent face**</mark>. SeaArt offers three methods to achieve this: <mark style="background-color:yellow;">**LoRA Training**</mark><mark style="background-color:yellow;">,</mark> <mark style="background-color:yellow;"></mark><mark style="background-color:yellow;">**Face Swap**</mark><mark style="background-color:yellow;">,</mark> and <mark style="background-color:yellow;">**Identical Seed**</mark>. You can choose **one of these methods** for character creation.

#### **LoRA Training**

Create <mark style="background-color:yellow;">**a dataset of the same face**</mark> (with about 50 images) and <mark style="background-color:yellow;">remove all facial feature tags</mark> to achieve a consistent appearance. For the base model, we suggest \<unStable Illusion> for European faces and \<Beautiful Realistic Asians> for Asian faces.

When generating images with LoRA, it's recommended to use <mark style="background-color:yellow;">**the base model from the training**</mark> (European: unStable Illusion, Asian: Beautiful Realistic Asians). The suggested LoRA weight is between 0.6-1. Please pay attention to maintaining <mark style="background-color:yellow;">**consistent LoRA weights and base models**</mark><mark style="background-color:yellow;">.</mark>

<mark style="background-color:yellow;">**Prompts (keep them simple)**</mark><mark style="background-color:yellow;">:</mark> Include character traits such as age, hair color, and eye color. Add other info as needed, such as clothing and scenery. (For example: 1girl, Brown hair, Realistic, Working out in a gym.)

If the face generated is not ideal, use Img2Img -> **Character Repair** with a Denoising Strength of around 0.3 (using the previous large model + LoRA + prompts).

#### **AI Face Swap**

<figure><img src="/files/fBuNMCjSR5Lz93dz28TW" alt="AI Face Swap" width="563"><figcaption><p>Before / After</p></figcaption></figure>

**Face Swap Template Sources:**

1. SeaArt's **built-in image/video templates**
2. Generate scenes and poses, then swap the face
3. Stock image/video websites

#### **Identical Seed**

After generating an image with a large model, you can **copy the seed** and use it with the same model, along with **slightly adjusted prompts**, to achieve a consistent facial appearance.

### **Publish Content & Generate Copy**

Once the image is finalized, you can go on to set up an account.

**Personal Profile**

Welcome, everyone, I'm xxx / your xxx,

I'm from xxx / a Berlin girl

Personality: Funny, lovable, likes xxx

Other Links...

<figure><img src="/files/vKUES5VuPe8o9PmXFb3U" alt=""><figcaption></figcaption></figure>

**Update Content**

**a. Generate Copy**

Utilize SeaArt's **AI Chat (or other large language models)**, select **Social Media KOL**, and automatically generate copy.

<figure><img src="/files/Ond4h1EVZiRcrAtd3zGl" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/s1hjALJW8k7s6Pcly2x6" alt="Generate Copy" width="563"><figcaption></figcaption></figure>

It could be about everyday life (fitness, swimming, shopping...), edgy stuff, greetings, etc. Recommended Tags: #aigirlfriend #aigirl #aimodel #ailingerie #aifashion#virtualmodel #digitalgirl #aiinfluencer #seaartai

**b. Video Content**

You can upload your favorite templates (or use ones provided by SeaArt) and use [**SeaArt.AI**](http://SeaArt.AI)**'s Face Swap feature** to swap in the virtual influencer's face before posting them on social media platforms.

After setting up an account, you can consistently update social media content by following this procedure: **Generate Images** -> **Update Content** -> **Publish Content**.

<mark style="background-color:yellow;">**By using this method, you can achieve automated operations, quickly create multiple virtual influencer IPs, and operate several matrix self-media accounts on your own, thereby monetizing through various AI influencer IPs.**</mark>

With the continuous development of AI technology, using the IP of virtual influencers for content monetization has become a trend in Internet marketing. If you have already planned to earn significant profits through this approach, this article will help you start. Please note that managing social media accounts requires persistence. Follow the steps outlined above to create and regularly operate AI influencers, and you are bound to attract considerable traffic and fans, leading to substantial profits.


# LOGO Design

Learn about the logo styles and use AI to generate your brand's perfect logo with these tips and tricks.

> Logos are generally divided into 4 styles:

**Brandmark Logo:** Uses images or icons to represent the brand intuitively.

<figure><img src="/files/DQ3tNK354FgWC4z9Za8i" alt="Brandmark Logo"><figcaption></figcaption></figure>

**Lettermark Logo:** Uses initials or abbreviations to represent the name of the company or brand.

<figure><img src="/files/Nm2eUV7j9n8aGVSBJRs7" alt="Lettermark Logo"><figcaption></figcaption></figure>

**Combination Mark Logo:** Combines words with images. They can be used separately or together.

<figure><img src="/files/TUR59dOYQwojkRpMhUhi" alt="Combination Mark Logo"><figcaption></figcaption></figure>

**Emblem Logo:** Often used by organizations such as automobile and sports teams. This style is more traditional.

<figure><img src="/files/ilJAGSkmLJYo6E9Hbuz2" alt="Emblem Logo"><figcaption></figcaption></figure>

We can consider the logo content based on our needs and enter the prompts.

Please make your prompts as clear as possible with all logo elements including the main body, colors, and style, The overall process includes: **image quality**+ **logo main body**+ **background**+ **style** and **some details**.

For example, we're designing an animal-type logo.

> Open [**SeaArt.AI**](http://seaart.ai/) and select the **SDXL** model.

<figure><img src="/files/DJvyqI5DJC25zHFnpLV6" alt=""><figcaption></figcaption></figure>

> **Parameter settings:**

**Model**: SD XL

**LoRA:**

1.Logo.Redmond - Logo Lora for SD XL 1.0 (0.85)

2.Texta - Generate text with SDXL (0.65)

**Prompts:**

logo\_l, fox, white background, simple background, fire, animal ears, closed eyes

**Negative Prompts:**

bad art, ugly, deformed, watermark, duplicated

**Image Size:** 1024x1024\[1: 1]

After the generation is completed, you can select your favorite image to further improve it.

Use Photoshop to **adjust details, edit fonts, or optimize the layout**, Make sure the logo colors correspond to your brand color scheme.

<figure><img src="/files/bD2ZCz13M58tQ4xhXQ0t" alt="AI Designed Logo"><figcaption></figcaption></figure>

We've prepared a series of prompts that are frequently used to create logos, You can enter different prompts based on your needs for the logo styles.

> **Image quality:** best quality, masterpiece, super detail, ccurate, high quality, highres, 4K, 8K

> **LOGO types:** logo, icon, badge…

> **Background:** light background, white background, simple backgroun

> **Style:** 3D, minimalistic, techy, outlined, hand-drawn, luxury, medieval, antique, stroke, processual, pixel

> **Details:** with the letters, with square, geometric shapes(rectangle, square, triangle, circle), no text


# E-commerce Poster

Design stunning e-commerce posters effortlessly! Leverage SeaArt Canvas's AI-powered tools to create captivating visuals and elevate your brand.

> SeaArt Canvas is divided into Realtime Canvas and Generate Canvas.&#x20;
>
> In **Realtime Canvas**, we can view the image results in real-time and make adjustments to address the uncontrollable issues of AI image generation.&#x20;
>
> In **Generate Canvas**, we can directly use the image generation function on the canvas, enabling convenient use AI functions such as partial repaint and erasure on images.

SeaArt Canvas offers deep involvement in the design process, with highly controllable AI drawing capabilities that turn your imagination into reality. Next, let's experience the functionality of SeaArt Canvas through a practical e-commerce poster example.

## **Product Background**

<figure><img src="/files/vvuAXzaRn5hdBHVomC9B" alt="Product Background" width="188"><figcaption></figcaption></figure>

【Brand New Upgrade · Ultimate Experience】The XS4 is a meticulously crafted gaming controller, utilizing the latest technology and premium materials to bring you an unprecedented gaming experience. Its ergonomic design ensures comfortable grip even during extended gaming sessions, while responsive buttons and precise joysticks guarantee accurate operation with every input.

&#x20;【Wide Compatibility】Whether it's PC, PS4, Xbox, or Switch, this multi-platform gaming controller is perfectly compatible. With no complicated setup required, it's plug-and-play, allowing you to enjoy gaming fun anytime, anywhere.

&#x20;【Long-lasting Durability】With a built-in high-quality battery, it supports up to 30 hours of continuous gameplay, ensuring uninterrupted gaming sessions. Its highly wear-resistant surface treatment and robust structure ensure the controller can withstand daily wear and tear.&#x20;

【Professional-grade Performance】Whether for casual entertainment or competitive gaming, this controller delivers top-notch performance. Enhanced vibration feedback and programmable buttons add more fun and strategic depth to your gaming experience.

## **Preliminary Consideration**

**I.** Because the controller belongs to the gaming and technology product category, this can lead to concepts such as science fiction, high technology, and alien worlds. With these directions in mind, we can use AI to generate images and add inspirational materials. Open the SeaArt Canvas, select <mark style="background-color:yellow;">**Generate Canvas**</mark><mark style="background-color:yellow;">,</mark> enter the prompts in the image generation box, choose <mark style="background-color:yellow;">**4**</mark> for the number of images, and leave other options at default parameters. Click on Generate Images to obtain the generated pictures.

<figure><img src="/files/HPmRIUB7pgn3lw3yLE0z" alt="SeaArt AI Canva Interface"><figcaption></figcaption></figure>

**II.** In this mode, the generated images will remain on the canvas, allowing you to continue generating images by moving the image generation box aside, much like drawing on a sketchpad. You can delete unnecessary images and keep useful ones as references to more intuitively adjust image parameters.

<figure><img src="/files/uIjDncgtciyhEUI4j8Mb" alt="SeaArt AI Canva - Adjust Image Parameters"><figcaption></figcaption></figure>

**III.** Once you have selected the reference images, we can proceed to the next step, which is to use the <mark style="background-color:yellow;">**Realtime Canvas**</mark> to add backgrounds to the product.

<figure><img src="/files/lZ7Wq6ias3GjDyTOGZ5f" alt="SeaArt AI Realtime Canvas" width="563"><figcaption></figcaption></figure>

## **Realtime Canvas**

### **Step1: Upload the Product**

Click on the <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Product**</mark><mark style="background-color:yellow;">"</mark> option on the left side, and then select the <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Upload Product**</mark><mark style="background-color:yellow;">"</mark> function. After uploading the product image, the system will automatically perform the clipping operation to ensure that the product image remains clear and accurate during the drawing process.

<figure><img src="/files/POV9cmGQQtlR9zQORyUq" alt="Realtime Canvas - Upload Product"><figcaption></figcaption></figure>

The scale adjustment tool is provided in the top left corner of the canvas, with a maximum output resolution of up to 1024x1024 pixels. Below that is the prompt box, where we can enter the prompts from the previous reference images.

<figure><img src="/files/TQPTISBOWoQN3wQCU5vp" alt="Adjust Output Resolution" width="563"><figcaption></figcaption></figure>

### **Step2: Generate Background**

Next, let's start generating the background for the product. Select the <mark style="background-color:yellow;">**simple Pencil**</mark> tool from the toolbar, and use the dropdown options to adjust the brush size and color. The layer button at the bottom left allows users to adjust the order of each layer, flexibly managing the overlap between layers.

<figure><img src="/files/y5exh8HokGr9nCSJx0NK" alt="Realtime Canvas - Simple Pencil"><figcaption></figcaption></figure>

Before starting the generation process, adjust the relevant parameters such as <mark style="background-color:yellow;">**Denoising strength (recommended to set to 0.7), seed (-1)**</mark>, etc., to optimize the generation effect. Then, click on the <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Started/Stopped**</mark><mark style="background-color:yellow;">"</mark> button to begin the generation.

Adjust the final effect of the image by increasing or decreasing strokes and prompts. Once you have obtained the desired image, click on the download button in the top right corner to save the picture.

<figure><img src="/files/EaBHZ2GmsbFvLOENh0t5" alt="Adjust the Final Effect of the Image"><figcaption></figcaption></figure>

**Merge Background**

In addition to drawing the background manually, we can also directly use the previous reference images as the background. The SeaArt Canvas can seamlessly integrate the product with the background.

In <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Generate Canvas**</mark><mark style="background-color:yellow;">,"</mark> right-click on the selected image and choose <mark style="background-color:yellow;">"</mark><mark style="background-color:yellow;">**Save to Mine.**</mark><mark style="background-color:yellow;">”</mark>

<figure><img src="/files/3wkMo78fCx8Ophz21e3u" alt="Save to Mine"><figcaption></figcaption></figure>

Switch to <mark style="background-color:yellow;">**"Realtime Canvas,"**</mark> click on "Mine," select the saved image, and add it to the canvas.

At this point, you can <mark style="background-color:yellow;">**reduce the Denoising strength**</mark> to preserve more details of the original background image. As you can see, through the canvas, we can seamlessly integrate the product into the background, creating a natural and harmonious visual effect.

<figure><img src="/files/xI5IVh0XEcFimn3yXIqr" alt="Reduce the Denoising Strength"><figcaption></figcaption></figure>

### **Step3: Enhancing Image Quality**

Due to the size limitation of the canvas, we also need to enhance the image quality of the final picture.

Open the <mark style="background-color:yellow;">**Swift AI**</mark> and select the <mark style="background-color:yellow;">**AI Image Upscaler.**</mark>

Import the image that needs enhancement and adjust the specific parameters according to your requirements.

<figure><img src="/files/XL6GAAKXywKmQfLN0Zcx" alt="Adjust Parameters"><figcaption></figcaption></figure>

### **Step4: Final Result**

Finally, edit the enhanced image and add necessary text descriptions. Our e-commerce poster is now complete.

<figure><img src="/files/nGfFsVzXiTjz6FhPqdM3" alt="E-commerce Poster by SeaArt AI "><figcaption></figcaption></figure>

With the continuous advancement and innovation of AI technology, future versions of SeaArt Canvas will introduce more intelligent features, such as deeper image analysis and generation algorithms. This will further streamline the design process and enhance creative efficiency.


# How To Make Multiple Characters

This guide will teach you about all details, useful tools and tricks to create perfect couple images including suggested models, LoRAs and effective prompting.

## General Tips | Prompting Tricks

● Use <mark style="background-color:yellow;">couple</mark> tag before your subjects and add <mark style="background-color:yellow;">hetero-yaoi-yuri</mark> depending on what type couple you want to generate.

● Using <mark style="background-color:yellow;">duo focus</mark> is effective at avoiding pov views and generating both subjects at same image.

● More characteristic term in prompt cause more randomized results, try to choose model that recognize your character or use LoRA for it and keep characteristics simple. Avoid to use many terms like <mark style="background-color:yellow;">black hair, green eyes, blue hair</mark> this type terms can effect both of subjects or wrong one without <mark style="background-color:yellow;">BREAK</mark>, but remind that BREAK is not 100% working technique.

● Use 2boys instead of 1boy, if you are going to make <mark style="background-color:yellow;">yaoi(gay)</mark> couple. | Use <mark style="background-color:yellow;">2girls</mark> instead of <mark style="background-color:yellow;">1girl</mark>, if you are going to make <mark style="background-color:yellow;">yuri</mark>(lesbian) couple.

● Use d<mark style="background-color:yellow;">uplicate, bilateral symmetry</mark> in negative prompt to avoid similar characters.

### Basic Examples

● Let's say i want to create Mario x Ahri couple, i must start write like this for my subject: *couple, hetero, duo focus, 1girl, 1boy*

● If i want to create Makima x Sakura couple: *couple, yuri, duo focus, 2girls*

**Note:** Prompting can show difference depending on model you use, this is only an example based on Danbooru tags and Model's general training.

## Creating couple Images With SD1.5 Models

● Most of anime models has pretty small character memory because of their training.

● Generally models recognize characters from most popular Danbooru Character Tags and also some of quite popular characters such as Mario, certain league of legends characters, protagonist of popular games etc.

● For detailed and accurate couples, LoRAs become important part of generation.

● <mark style="background-color:yellow;">BREAK</mark> keyword is effective for keeping character's characteristics seperate and don't blend them.

### Example SD1.5 Nami X Naruto

**Positive Prompt**

> (masterpiece, best quality:1.2), highly detailed, (couple), hetero, (duo focus), 1girl, 1boy, holding hands, looking at another, beach scenery, bokeh, (depth of field:1.2), absurdres, highres, soft lighting, soft shadows, BREAK, 1girl, nami (one piece), NamiFinal, ginger, long hair, pale skin, busty, bikini, BREAK, 1boy, uzumaki naruto, blonde, short hair, shorts

**Negative Prompt**

> EasyNegativeV2, (badhandv4), nsfw, duplicate, (blending), (bilateral symmetry:1.5)

**LoRA**

[Nami (ナミ) One Piece Character LoRA (Post-timeskip) ](https://www.seaart.ai/models/detail/7aaf49dd97da7a7d048761dfbc76e7bd)

[Naruto Uzumaki / 漩涡鸣人 / ずまき ナルト / 火影忍者](https://www.seaart.ai/models/detail/680d8d4a5bf576396fa1e63990c91368)

**Embedding**

* \[TI] EasyNegativeV2 \[Textual Inversion Embedding]
* badhandv4

**Model** ⁠[AnythingElse](https://www.seaart.ai/models/detail/589f4792f55f6de8c6e4b4d6630e5082)

**Explanation**&#x20;

As you can see i used `BREAK` here and made 3 seperate prompt to avoid blending.

1. Part: General prompt, you must use both subjects here otherwise models will create single subject. I used `(couple), hetero, (duo focus), 1girl, 1boy` for subjects and add my additional prompts.
2. Part: Name of first subject and details, i used `1girl` for first subject and add her characteristics and LoRA to make her look like Nami from One Piece.
3. Part: Name of second subject and details, i used `1boy` for second subject and add his characteristics and LoRA to make him look like Uzumaki Naruto from Naruto.

### Thoughts About SD1.5

* You can create ship images with LoRAs and `BREAK` but i have to admit it, it's hard and not worthy.
* Most of SD1.5 Models towards to generate single subject and you can't even get 2 person in same image.
* Prompt understanding is not that strong, even with `BREAK` there is a still chance to get mixed characteristics. `BREAK` is not 100% working trick.
* There is high change to get bad and random compositions.

## Creating Ship Images with SDXL Models

* They have bigger character database compared to SD1.5 Models and creates more accurate characters.
* Generally models recognize most of characters from Danbooru Characters Tags and many well-known characters from different genres such as movie, anime, game etc.
* LoRAs are not necessary but might be useful at some creations.
* **Animagine XL** is very great choice to make ship images, it's better than other SDXL based models and has larger character database.

### Example SDXL Boa X Luffy

**Positive Prompt**&#x20;

> couple, duo focus, hetero, 1girl, 1boy, boa hancock hugging to Luffy, one piece, ship background, lovely, masterpiece, best quality, safe, SFW, very aesthetic, recent, absurdres, highres

**Negative Prompt**&#x20;

> (worst quality, low quality:1.3), NSFW, explicit, oldest, early, (very displeasing, displeasing), lowres, bad anatomy, anatomical nonsense, bad hands, jpeg artifacts, ((ugly)), grayscale, traditional media, duplicate

**Model** [⁠Animagine XL](https://www.seaart.ai/models/detail/f2755cd95dd840080d622ca62e381fc8)

### Thoughts About SDXL

It's overall good, especially Animagine XL model is so good at recognizing characters, this means you don't have to prompt details and there will be no blended characteristics most of time.

## Creating Ship Images With Pony Models

* Pony models are actually SDXL based and this means they are atleast good option as SDXL models.
* They are very customizeable so you can create almost every type couple you can imagine.
* Generally recognizes most of well-known characters from different genres including anime, game, cartoon, comics etc.
* **AutismMix** is good for anime couples while **T-Ponynai3** especially excels at Genshin Impact characters.

### Example Pony Hu Tao X Eula

**Positive Prompt**

> score\_9, score\_8\_up, score\_7\_up, score\_6\_up, score\_5\_up, score\_4\_up, (anime), sfw, couple, yuri, duo focus, 2girls, hu tao (genshin impact), eula (genshin impact), carrying another on arms, hugging, lovely couple, long wedding dress, wedding suit, wedding ceremony, outdoors, (nature)

**Negative Prompt**

> (score\_4, score\_3, score\_2, score\_1), ugly, bad anatomy, bad hands, nsfw, duplicate, blending

**Model** ⁠[T-Ponynai3 V6 ](https://www.seaart.ai/models/detail/34ec36d43cf42880bf985f83fe3d7d85)

### Thoughts About Pony

* It's overall good and has various options for different concepts.
* Characteristics can blend sometimes depending on what you try to generate.
* Creative and impressive results.
* Good at rendering details


# Prompt Templates

Here are some commonly used prompt templates. Simply change parts of the phrases, and you can generate equally high-quality images.

Don't know how to write prompts? Try these prompt templates; you only need to modify elements such as the **subject, clothing, and time.**

## Desert Wasteland

Masterpiece, **\[angle]**, transitioning to a desert landscape at **\[time of day]**, **\[subject]**, wearing **\[clothes]**. Layers of tattered fabric, unconventional accessories, and a weathered look create an aura of survival and resilience, 16K, ultra high resolution, photorealistic, UHD, RAW, DSLR, post-apocalyptic landscape with ruins and debris, Mad Max vibes.

## Moonlit scenery

**\[subject]** standing in a sunlight landscape, (best quality, 4k, 8k, highres, masterpiece:1.2), ultra-detailed, (realistic, photorealistic, photo-realistic:1.37), detailed facial features, **\[expression], \[clothes],** elegant pose, dramatic moonlit backdrop, mystical atmosphere, (fantasy, cinematic:1.2), ethereal lighting, cool toned color palette, chiaroscuro lighting.

## Anthropomorphic animals

<figure><img src="/files/IxdrC5ss7qCwOfBHAbsE" alt=""><figcaption></figcaption></figure>

A sophisticated anthropomorphic **\[animal]** dressed in a **\[clothes]. \[Fur],** and **\[appearance].** The scene is **\[environment]. \[time].**

## Fashion Magazine (Anime Character)

Model: Animagine XL V3.1

(masterpiece), detailed art, (fashion magazine cover:1.5), **\[1 boy/girl]**, **\[Character Name]** / (Genshin Impact), **\[Character Traits], \[Body Type], \[Weapon and Its Features], \[Atmosphere],** text says "Fashion Show", (cover), top quality, pop colors.


# How to Maintain Character Consistency

Maintaining character consistency is crucial in AI-generated art for high-quality results. Consistent features, outfits, and details enhance coherence and immersion.

In AI-generated art, maintaining character consistency is essential for creating high-quality character images. Consistent appearance, outfits, and details enhance the coherence of the work and improve audience immersion. This guide introduces two methods to ensure the character’s appearance remains consistent across multiple generations.

## Method 1: Fixed Seed

This is the simplest and fastest way, often used to create AI influencers.&#x20;

The steps are as follows:

**Model**: The Most Pure Realism

**Step 1**

Generate a random character image and obtain the seed from the image.

**Seed**: 397096412

> Prompts: Retro, beautiful female model, attractive figure, long hair, blonde, smooth skin, blue eyes, delicate makeup, beige lips, heavy eye makeup, ((draped suspender dress)), cardigan, knitted sandals, realism, open air, hyperrealistic 8K, ultra-detailed realistic, pose sculpture, hair blowing in the wind

**Step 2**

Use the same seed as the previous image (fixed seed) and modify some of the prompts. You can freely change the character’s background, outfit, pose, etc., while keeping other parameters the same.

> Prompts: Retro, beautiful female model, seductive figure, long hair, blonde, smooth skin, blue eyes, delicate makeup, beige lips, heavy eye makeup, <mark style="background-color:yellow;">wearing yoga pants,</mark> realism, <mark style="background-color:yellow;">grass,</mark> hyperrealistic 8K, ultra-detailed lifelike, hair blowing in the wind

> Prompts: Vintage, beautiful female model, seductive figure, long hair, blonde, smooth skin, blue eyes, delicate make-up, beige lips, heavy eye makeup, <mark style="background-color:yellow;">wearing a white dress,</mark> realistic, <mark style="background-color:yellow;">on the beach, she wears a shell necklace around her neck and braided bracelets on her wrists,</mark> hyperrealistic 8K, ultra detailed realistic

<figure><img src="/files/iJLHGvMCvZEZ513R7Dam" alt="" width="375"><figcaption></figcaption></figure>

> Prompts: Retro, beautiful female model with seductive figure, long hair, blonde, smooth skin, blue eyes, delicate make-up, beige lips, heavy eye makeup, <mark style="background-color:yellow;">wearing a thick down jacket with matching scarf and gloves,</mark> photo-realistic, <mark style="background-color:yellow;">with a fur hat on her head, climbing on snowy mountains,</mark> surreal 8K, ultra detailed realistic

This method is suitable when no reference image is available. The reference image is randomly generated by the AI, and subsequent images can retain the appearance of the first one.

## Method 2: Training LoRA

Train a character LoRA using reference images of the character and set a trigger word. Then proceed with image generation.

You can refer to the following article for a detailed LoRA training guide.

{% content-ref url="/pages/bDLwl9DD9y2ga18oYqz4" %}
[3-2  LoRA Training (Advance)](/guide-1/3-advanced-guide/3-2-lora-training-advance)
{% endcontent-ref %}

This method is suitable when you have a large collection of character images, resulting in more diverse outputs.

Both methods are applicable to real-life or anime characters.


# Useful  Prompts

This guide focuses on key terms and their meanings, enhancing outputs across various models, with the best performance on SDXL models, especially Juggernaut XL.

This guide focuses on key terms and their meanings for SeaArt users, especially those working in foreign languages. These keywords enhance outputs across various base models but perform best with SDXL models, particularly Juggernaut XL.

## About Subject

#### portrait

A formal photograph with style and settings, it's realistic but far from daily life.

#### candid portrait

A photograph captured without creating a posed appearance, it's better option for generating life-like renders that looks like part of reality. For example, *<mark style="background-color:yellow;">candid portraits, a woman drinking her coffee while watching outside</mark>*

#### animal portrait

A keyword that useful for creating animal focused images.

#### automotive photography

A keyword that useful for creating car, automobile focused images.

#### fashion photography

A keyword that helps to create more stylized people renders.

#### nature photography

A keyword that helps to create more natural results with green toned environment.

#### landscape/scenery

This keywords tends to create mesmerizing views from world or your fantasy, useful for generating amazing ambients.

## About Age

#### youthful

Creates young adult people without any childish detail. This keyword is useful for generating 20 years old likely people. `young` keyword can create illegal images especially on realistic models, prefer `youthful` instead of this.

#### mature

Creates grown-up adult people, this can be used for middle-aged human illustrations.

#### adult female/adult male

Adult keyword can be used with gender to avoid childish generations. It's especially good with non-human (furry, humanoid, animal) subjects, you can use this prompt instead of woman or girl because they will make AI focus on humans.

#### child, childish

Use this keywords in **negative** to avoid generating any illegal content. Illegal contents may end up with banning your account. Use "loli" keyword in negative for anime models with this keywords.

## About View, Angle

#### close-up

Very near view of subject.

#### ........ focus

hand, face like terms can be used with focus to create renders of that spesific area.

#### ........ shot

close, portrait, medium, long, far like terms can be used with shot to create renders with that spesific distance.

#### cowboy shot

A shot framed from the actor's mid-waist to right above their head.

#### dutch angle

A shot in which the camera has been rotated around the axis of the lens and relative to the horizon or vertical lines in the shot.

#### fisheye view

a wide-angle photographic lens that has a highly curved protruding front, that covers an angle of about 180 degrees, and that gives a circular image.

#### aerial view/bird's eye view

A viewpoint seen at a high elevation.

#### worm's eye view

A worm's-eye view is a description of the view of a scene from below that a worm might have if it could see. It is the opposite of a bird's-eye view.

#### ........ angle

Low, wide, high like terms can be used with angle to create images from that spesific angle.

#### ........ view

Front, back, side like terms can be used with view to create images from that spesific view.

### perspective

Perspective has several different meanings—several applicable in some way to photography.

#### linear perspective

A system of creating an illusion of depth on a flat surface. All parallel lines in a painting or drawing using this system converge in a single vanishing point on the composition's horizon line.

#### aerial perspective

Aerial perspective is created by weather conditions. Looking out over a landscape from a high vantage point.

#### forced perspective

This type of perspective involves manipulating the size and distance of objects to create a sense of scale and depth.

#### diminishing scale perspective

Diminishing scale perspective utilizes how we naturally view things with our eyes. Closer options appear larger and farther options appear smaller.

#### geometric perspective

Geometric perspective is a drawing method by which it is possible to depict a three-dimensional form as a two-dimensional image that closely resembles the scene as visualized by the human eye.

## About Color And Tones

#### high contrast

A high-contrast image will have very bright highlights and very dark shadows.

#### low contrast

Low contrast images have neither very deep shadows nor strong highlights which could direct the viewer's eye to a particular detail.

#### HDR/high dynamic range

A technique that expresses details in content in both very bright and very dark scenes.

#### monochromatic color palette

Monochromatic color schemes use a single color with varying shades and tints to produce a consistent look and feel. Although it lacks color contrast, it’s often very clean and polished. It also allows you to easily change the darkness and lightness of your colors.

#### analogous color palette

Analogous color schemes are formed by pairing one main color with the two colors directly next to it on the color wheel.

#### complementary color palette

Complementary color scheme is based on two colors directly across from each other on the color wheel and the relevant tints of those colors, provides the greatest amount of color contrast.

#### neutral color palette/neutral colors

Any group of colors that have been muted or desaturated.

#### warm color palette/warm color tone/warm colors

Warm colors include red, orange, and yellow, and variations of those three colors.

#### cool color palette/cool color tone/cool colors

Cool colors are typified by blue, green, and purple and their hybrids.

#### muted colors

Muted colors, also referred to as desaturated or subdued colors, are ones that have been softened by adding gray or a complementary color to reduce their brightness intensity.

#### faded colors

Having lost freshness or brilliance of color.

#### ........ theme

Red, blue, brown, aquamarine, crimson like terms can be used with theme to allow AI to prioritize that spesific colors in art.

#### colorful

This term can be used to have much and various colors in subject.

## About Quality

#### ........ skin textures

Realistic, natural like terms can be used with this prompt to increase reality of subject and decrease doll-like, cartoonish look of skin.

#### cinematic

Perfect for creating dynamic, narrative-rich images with a film-like quality.

#### high resolution/high-resolution

One of the most effective keywords that improve quality and details of image.

#### bad hands

This keyword can be used in **negative** to improve accuracy of hands.

#### 5-funny-looking-fingers

Most effective **negative** keyword to improve hands overall in images. It can be even more effective than certain LoRAs.


# 6-Permanent Events

Discover SeaArt events and activities.


# SeaArt.AI Creator Incentive Program

Earn real cash rewards for your AI creations! Join the SeaArt.AI Creator Incentive Program and get rewarded for your talent.

🙋‍<mark style="background-color:yellow;">**Join Quickly to Share the Bonus**</mark>

### What is the SeaArt Creator Incentive Program?

The SeaArt Creator Incentive Program (hereinafter referred to as the "Incentive Program") is a growth support initiative designed specifically for original creators. Through a three-part privilege system comprising cash incentives, identity honors, and official resource support, it helps creators unlock their potential and jointly build a diverse, high-quality AI content ecosystem.

### **Three Incentive Categories**

Original Models | AI Apps | AI Characters

Content Requirements:

Originality: Works must be 100% original.

Distinctiveness: Must demonstrate significant uniqueness in style, details, functionality, or visual effects.

Positive Value: Content must align with a positive and healthy creative direction.

### **Detailed Reward Privileges**

**😍Cash Incentives**

1. Cash Incentives

Independent Bonus Pools: Dedicated bonus pools for models, AI apps, and AI characters, with earnings directly linked to work contribution value.

Incentive Upgrade: <mark style="color:red;">Starting March 2025</mark>, total bonus pools across all categories will <mark style="color:red;">increase by 50% or more</mark>, benefiting more outstanding creators.

Dynamic Strategy: Our platform will continuously optimize incentive rules based on content ecosystem development trends (such as vertical domain-specific support) to help you grow effectively.

<mark style="color:red;">\*New Logic: The closer the content publication date is to the earnings calculation date, the higher the earnings weight coefficient, resulting in higher earnings.</mark>

2. SeaArt.AI Original Creator Badge

Identity Mark: The exclusive badge will be displayed on your homepage, showcasing your original creative strength and platform-certified identity.

3. Official Resource Support

Visibility Boost: Quality works may be selected for homepage recommendations, special features, and other prime exposure positions to attract fans quickly.

In-depth Cooperation: Opportunities for exclusive contracts and first-release partnerships to maximize commercial value and long-term returns.

Full-on Promotion: Utilize on-platform resources (banners, notifications) and off-platform co-branded exposure (social media, industry events) to build the creator's IP.

Priority Partnership: Gain priority access to official creative collaboration opportunities, technical beta testing, and other exclusive resources.

### **How to Participate:**

<https://www.seaart.ai/creationCenter/home>

<figure><img src="/files/G5MIP1N2F1zXtENu1SAQ" alt=""><figcaption></figcaption></figure>

🙌<mark style="background-color:yellow;">No Entry Barriers</mark>, Start Earning in <mark style="background-color:yellow;">3 Steps</mark>

Log in to your account and go to your Personal Center -> Click on \[Incentive Program]

Read and agree to the "Incentive Program Participation Rules"

### **Earnings Calculation Rules**

1. Earnings Calculation Logic

Earnings are based on the user value contribution of works, calculated dynamically using "user appreciation" (number of downloads, number of uses) and "user satisfaction" (share rate, number of derivative creations).

2. Calculation Cycle

Data Display: The previous month's earnings data is updated before the 5th of each month (Navigate to Incentive Program -> Data Dashboard).

Automatic Calculation: Creators' earnings are calculated monthly and automatically sent to their account the following month.

Withdrawal Rules: After meeting conditions, withdrawal can be requested from the 1st to the 15th of each month. Our platform will complete payment within 5 working days.

3. Special Notes

Original Protection: Only original content is eligible for earnings distribution; copyright-infringing content will result in disqualification and earnings recovery.

Fairness Guarantee: Strictly prohibited behaviors include fake engagement, puppet accounts; violators will be banned from participation.

<mark style="background-color:red;">Any other questions? Click here!👇</mark>

{% content-ref url="/pages/1lkFWefHqi78e0eYQeMZ" %}
[Creator Incentive Program FAQ](/guide-1/6-permanent-events/seaart.ai-creator-incentive-program/creator-incentive-program-faq)
{% endcontent-ref %}


# Creator Incentive Program FAQ

Find answers to common questions about the SeaArt.AI Creator Incentive Program, including earning potential and eligible content.

### **I. Earnings Related**

**Q1: Why am I not earning despite uploading content?**

A1: Please ensure uploaded content is original and model status is marked as "Original" to earn.

💡Note: If multiple base models and LoRAs from the same creator are used in one creation, the system will count it as using only one model by default.

**Q2: How can I earn more incentives?**

A2: SeaArt adjusts incentive focus each round based on user needs. You can follow platform announcements to learn about incentive focus adjustments and optimize uploaded content to ensure compliance with the latest incentive policies.

**Q3: When uploading a model, if prompted "model already uploaded," and if the model is returned to my account, will historical data be retained?**

A3: Historical data will be retained, but the calculation of earnings will start from the moment the model is claimed.

### **II. Withdrawal Related**

**Q4: Is there a minimum amount requirement for withdrawal?**

A4: A single withdrawal amount must not be less than $100. Earnings that do not reach this limit will accumulate to the next month.

**Q5: What should I do if my withdrawal request is rejected?**

A5: If your withdrawal request is rejected, please check these possible reasons:

Account information is incorrect or incomplete.

Abnormalities have been found during the platform review.

You can resubmit the withdrawal request after correction based on the rejection reason.

### **III. Data & Calculation Related**

**Q6: What should I do in case of earnings data update delays?**

A6: Earnings data typically has a statistical delay of 1-3 days. If it is not displayed after more than 3 days, please contact customer service for verification.

### **IV. Participation Rules Related**

**Q7: Can I participate in multiple incentive categories simultaneously?**

A7: Yes. Our system will automatically divide bonus pools into different categories, and you can participate in multiple incentive categories simultaneously.

**Q8: Can amateur creators participate?**

A8: All creators are welcome! Regardless of experience level, as long as your content meets the platform requirements, everyone has an equal opportunity for incentives.

**V. Content Optimization & Exposure Related**

**Q9: How can I attract attention from more users?**

A9: You can increase content exposure and attention through the following methods:

1. Optimize Content Performance

The ranking of platform content is calculated based on a combination of the following factors: number of uses, number of downloads, number of favorites, number of likes, and number of comments.

Optimization Suggestions:

Enhance the appeal of your content cover and description.

Maintain platform activity by updating content regularly.

Promote your original content on external platforms.

2. Apply for Official Recommendation

Click this link to apply for a recommendation.👇

<https://forms.gle/mhCq23akQUuu8QU78>


# High-Quality Models Recommendation

Use a high-quality model and generate stunning images with just a simple prompt, all in one click.


# SeaArt Infinity

SeaArt Infinity produces ultra-high-resolution images with accurate anatomy and minimal deformations, excelling in rendering precise text. This guide covers its key features and capabilities.

SeaArt Infinity has the highest potential among AI image generators. It can produce ultra-high resolution images in almost any art style and offers some of the best anatomically correct renders, minimizing body and hand deformations. Additionally, it excels at rendering accurate text on images. This guide will cover all the differences and capabilities of the SeaArt Infinity Model.

**Try the Model:** [SeaArt Infinity](https://www.seaart.ai/models/detail/f8172af6747ec762bcf847bd60fdf7cd)

## Before Start | Noticeable Differences

SeaArt Infinity ignores negative prompts, meaning it doesn't matter if the negative prompt box is present or not—it will be ignored anyway. It uses a different encoder system (T5) than other AI models, which introduces some differences in prompting:

1. Keyword Weighting does not work. e.g. `(term:1.3), [[term]], (((term)))`
2. Keyword Blending/Edit tools does not work.
3. BREAK does not effect.
4. Better understanding of natural language prompts.
5. It can create different subjects in image without any effort.

e.g. `dog and cat` generates two different anatomically correct animal on Infinity while most of other models tends to generate creatures that are a combination of these two animals.

<figure><img src="/files/ApqpRuXLu4UyCzag07kP" alt=""><figcaption></figcaption></figure>

## What This Model Excels At

* Generating high quality images without any "Quality Modifier" prompt.
* Generating desired images with correct details and placements even with confusing prompts cause problem on other models. e.g. `blue cube under red ball`
* Generating accurate text renders in images. It can struggle with very long texts but capability of model is enough to satisfy most of users.
* Generating almost any subject with correct prompting.
* Crafting images with any art style or medium instead of focusing one of them.
* Crafting ultra photorealistic images that are indistinguishable from real photographs.
* Generating correct anatomy and hand renders.
* Crafting images with SFW/NSFW content.
* Recognizes Art/Artist styles from references.
* Works great with both of natural language and tag based prompts. Long descriptive prompts are recommended for detailed compositions.

<figure><img src="/files/65MBuetSs3ORj9H2eVvM" alt=""><figcaption></figcaption></figure>

## Cautions For Optimal Usage

* Don't use brackets in your prompt.
* All settings are already chosen, only thing user have to do is prompting.
* Describe style of image to avoid random art style images.
* Works good with base SDXL resolutions.
* It's not best model to generate NSFW, better to wait for LoRA support or new FLUX.1 based models.
* Text must be used inside of quotation marks `"`.

## Prompt Style

SeaArt Infinity has only positive prompt input and has better understanding of prompts than other image generators. It's best to use descriptive natural language prompts and reinforce it with certain tag style prompts.

### Text Render | Typography

SeaArt Infinity has best text renders compared to other AI image generators, user must write `text` inside of quotation marks `"` to seperate it from prompt. e.g.

> text written on a road sign "SeaArt"

<figure><img src="/files/GK6y4WacZSANcLSDafqR" alt="" width="188"><figcaption></figcaption></figure>

### Unwanted Stuff

SeaArt Infinity doesn't have negative input but prompts like `a road without any car` can be used in positive prompt to avoid unwanted stuff such as `car`. This is not an 100% working method for all prompts but useful trick.

### Recommended Prompt Order

1. Style/Medium
2. Subject
3. Subject Details
4. Additional Details
5. Lighting/Shadow
6. Additional Modifiers

#### Natural Language Prompt

Start with your subject and describe subject's details, style and medium. Then add additional details to create composition of image such as lighting, background, effects. After that you can add last modifiers to complete your prompt.&#x20;

**Note:** Dot `.` can be used inside of prompt after finishing sentences.

#### Suggested Usage

Use natural language prompt and don't seperate every word with comma `,` then reinforce composition of image with additional tags.

## Prompt Examples

**Prompt Understanding Ability**

<figure><img src="/files/UOWiWASDnaxbubtRkvFH" alt="" width="375"><figcaption></figcaption></figure>

**Positive Prompt**

> a pig hugging to a television from side, "UwU" written on television screen, a monkey sitting on top of television, monkey eating banana, octane render, rendered by unreal engine, highly realistic render, digital media artwork

**Photorealism**

<figure><img src="/files/08UaaJHiSaWP1nHWO2Ho" alt="" width="375"><figcaption></figcaption></figure>

**Positive Prompt**

> professional photographic shot of a beautiful woman, she has slightly wavy long ginger hair, her green eyes shining with reflection of light, she sitting on attic, sunshine comes inside from single square window, she smiling to the viewer, slightly tilted head, pastel bokeh effect, depth of field, subsurface light scattering, backlighting, dark ambient.

**Text Layout**

<figure><img src="/files/J06zJdFCZAQ5uiWuupCJ" alt="" width="375"><figcaption></figcaption></figure>

**Positive Prompt**

> first-person perspective, illustration of a hand holding letter, letter text "NOOT NOOT Joke aside it's really good at typography", pencil sketching, messy lines with hatching style, old yellowish paper, old fashioned fantasy theme.


# Stable Diffusion 3.5

Stable Diffusion 3.5 has impressive prompt understanding and adherence skills with its high-quality image creations. This guide will provide a detailed explanation of how to use SD3.5.

Stable Diffusion 3.5 is a new base model created by Stability AI with 8.1 billion parameters. The model has impressive prompt understanding and adherence skills, producing high-quality image creations. SD3.5 has great knowledge of mediums, styles, artists, mangakas, illustrators, and more, making it very flexible in its usage.

<figure><img src="/files/cV0izxfujSYbrbLvU3T9" alt=""><figcaption></figcaption></figure>

Try the Model: [Stable Diffusion 3.5](https://www.seaart.ai/models/detail/451c32af1283b24548d343a06d0bdb01)

## Before Start | Noticeable Differences

New Stable Diffusion has some similarities to Flux model in prompting, this means SD3.5 is also tend to natural language prompting and doesn't allow some features as Flux did before

1. Keyword Weighting does not work. e.g. `(term:1.3), [[term]], (((term)))`
2. Keyword Blending/Edit tools does not work.
3. BREAK does not effect.
4. Better understanding of natural language prompts.
5. It can create different subjects in image without any effort.

e.g. `dog and cat` generates two different anatomically correct animal on Stable Diffusion 3.5 while most of other models tends to generate creatures that are a combination of these two animals. However, Stable Diffusion 3.5 doesn't ignore negative prompts. Negative prompts are not necessary for model to work good but users still have option to use unlike the Flux. Also it has outstanding art, artist style knowledge compared to Flux, as many of you know Flux was pretty weak at understanding references and recognizing them but SD3.5 still have strong database of styles like its older versions. This makes SD3.5 great option for users couldn't find enough artist knowledge and beauty in Flux.

## What This Model Excels At

* Generating high quality images without any "Quality Modifier" prompt.
* Generating desired images with correct details and placements even with confusing prompts cause problem on other models. e.g. `blue cube under red ball`
* Generating accurate text renders in images. It can struggle with very long texts but capability of model is enough to satisfy most of users.
* Generating almost any subject with correct prompting.
* Crafting images with any art style or medium, it has great dataset of art styles.
* Crafting ultra photorealistic images that are indistinguishable from real photographs but Flux is better than SD3.5 in a terms of realism.
* Generating correct anatomy and hand renders but it still struggle with certain parts of body such as neck, chin, cheekbones etc.
* Crafting images with SFW content.
* Recognizes Art/Artist styles from references.
* Works great with both of natural language and tag based prompts. Long descriptive <mark style="background-color:yellow;">natural language prompts</mark> are recommended for detailed compositions.

## Cautions For Optimal Usage

* Don't use brackets in your prompt.
* You can use it from [AI App ](https://www.seaart.ai/workFlowAppDetail/cscbbqte878c7398hpmg)or standart generation page.
* Works good with base SDXL resolutions.
* It's not best model to generate NSFW, better to wait for LoRA support or new SD3.5 based models.
* Text must be used inside of quotation marks `"`.
* Negative prompts are not necessary but still can be used for unwanted stuff.

## Recommended Settings

* **Steps:** 20 (can be increased up to 40)
* **CFG Scale:** 4.5 (7 can be used for anime, cartoon like arts)
* **Sampling Method:** `euler`+`beta`, `euler`+`sgm_uniform`, `dpmpp_2s_ancestral`+`sgm_uniform`, `dpmpp_2m`+`sgm_uniform`

## Prompt Style

SD 3.5 has strong natural language prompt understanding capabilities with its ease of use style, model also supports negative prompt box so we can edit our images in easier and more effective way compared to Flux. It works good with tag like prompting but <mark style="background-color:yellow;">natural language prompts</mark> are recommended for best outputs.

### Text Render | Typography

Stable Diffusion 3.5 has outstanding text renders compared to other AI image generators, user must write `text` inside of quotation marks `"` to seperate it from prompt. e.g.

> text written on a road sign "SeaArt"

<figure><img src="/files/BuLsPsxrhDzXxACq7twR" alt="" width="188"><figcaption></figcaption></figure>

### Recommended Prompt Order

1. Style/Medium
2. Subject
3. Subject Details
4. Art/Artist References
5. Additional Details
6. Lighting/Shadow
7. Additional Modifiers

#### Natural Language Prompt

Start with your <mark style="background-color:yellow;">subject</mark> and describe <mark style="background-color:yellow;">subject's details, style and medium</mark>. Then add additional details to create composition of image such as <mark style="background-color:yellow;">lighting, background, effects.</mark> After that you can add last modifiers to complete your prompt.&#x20;

<mark style="color:red;">**Note:**</mark> Dot `.` can be used inside of prompt after finishing sentences.

#### Suggested Usage

Use natural <mark style="background-color:yellow;">language prompt</mark> and don't seperate every word with comma `,` then reinforce composition of image with additional tags.

## Prompt Examples

### Art Nouveau Portrait

<figure><img src="/files/fSK0ZNrAXldLav0MjP6U" alt="" width="563"><figcaption></figcaption></figure>

**Positive Prompt**

> A detailed portrait of a woman with flowing hair and floral motifs, in the style of Alphonse Mucha, art nouveau, intricate patterns, iridescent watercolor ink splashes, high contrast colorful artwork, brush lines over image that remind us confusion of life.

**Negative Prompt**

> nsfw, nudity, child, childish

### Winter Cabin

<figure><img src="/files/Ub9sjBdLvdp2NFrhz14b" alt="" width="563"><figcaption></figcaption></figure>

**Positive Prompt**

> Cinematic photography, A cozy winter cabin nestled in a snowy forest, soft golden light glowing from the windows, surrounded by tall pine trees. Realistic style, high detail, warm atmosphere, winter wonderland scene.

**Negative Prompt**

> blurry, dark, empty.

### Embrace The Darkness

<figure><img src="/files/WFbOfcrkLHJEPPL2OE4x" alt="" width="563"><figcaption></figcaption></figure>

**Positive Prompt**

> A dark, ominous medieval landscape with a lone warrior holding a large sword, in the style of Kentaro Miura, highly detailed, "Embrace The Darkness" written in handwrite on the sky, intense shadows, dark fantasy.

**Negative Prompt**

> nsfw


# SeaArt Realism

SeaArt Realism focuses on high-quality cinematic photography and realism, making it ideal for photography enthusiasts. Check the guide for more tips on using it.

SeaArt Realism is another Flux.dev based model from developers of SeaArt, Infinity was a versatile model that can produce many art styles but versatility reduces overall quality of styles. SeaArt Realism model trained to focus on cinematic photographs with its high quality standarts for realism. If you want to focus on photography, this model is tailored for you.

Try the Model: [SeaArt Realism](https://www.seaart.ai/models/detail/cb2801e3136f4d4c9c1c12dc852691c5)

## Before Start | Noticeable Differences

FLUX Models ignore negative prompts, which means it doesn't matter if Infinity has negative prompt box or not, it will be ignored anyway. It has different encoder system (T5) than other AI models and it bring some differences inside of prompting:

1. Keyword Weighting does not work. e.g. `(term:1.3), [[term]], (((term)))`
2. Keyword Blending/Edit tools does not work.
3. BREAK does not effect.
4. Better understanding of natural language prompts.
5. It can create different subjects in image without any effort.

e.g. `dog and cat` generates two different anatomically correct animal on Infinity while most of other models tends to generate creatures that are a combination of these two animals.

## What This Model Excels At

* Focused on realism and highly tends to maintain realism in overall composition.
* Crafting ultra photorealistic images that are indistinguishable from real photographs.
* Generating high quality images without any "Quality Modifier" prompt.
* Generating desired images with correct details and placements even with confusing prompts cause problem on other models. e.g. `blue cube under red ball`
* Generating accurate text renders in images. It can struggle with very long texts but capability of model is enough to satisfy most of users.
* Generating almost any subject with correct prompting.
* Generating correct anatomy and hand renders.
* Generating both genders with any ethnicities.
* Works great with both of natural language and tag based prompts. Long descriptive <mark style="background-color:yellow;">natural language prompts</mark> are recommended for detailed compositions.

## Cautions For Optimal Usage

* Don't use brackets in your prompt.
* All settings are already chosen, only thing user have to do is prompting.
* Don't use any art terms to confuse model and push it to break realism.
* Works good with base SDXL resolutions.
* It's not best model to generate NSFW, better to wait for LoRA support or new FLUX.1 based models.
* Text must be used inside of quotation marks `"`.

## Prompt Style

FLUX.1 has only positive prompt input and has better understanding of prompts than other image generators. It's best to use <mark style="background-color:yellow;">descriptive natural language prompts</mark> and reinforce it with certain tag style prompts.

### Text Render | Typography

FLUX.1 has best text renders compared to other AI image generators, user must write `text` inside of quotation marks `"` to seperate it from prompt. e.g.

```
text written on a road sign "SeaArt"
```

<figure><img src="/files/KKtFfwLs4qZYi4mkxx4M" alt="" width="172"><figcaption></figcaption></figure>

### Unwanted Stuff

Unfortunately FLUX.1 doesn't have negative input but prompts like `a road without any car` can be used in positive prompt to avoid unwanted stuff such as `car`. This is not an 100% working method for all prompts but useful trick.

### Recommended Prompt Order

1. Photography/Shot Type
2. Subject
3. Subject Details
4. Additional Details
5. Lighting/Shadow
6. Additional Modifiers

#### Natural Language Prompt

Start with your subject and describe subject's details and how it looks. Then add additional details to create composition of image such as lighting, background, effects. After that you can add last modifiers to complete your prompt.&#x20;

<mark style="color:red;">**Note:**</mark> Dot `.` can be used inside of prompt after finishing sentences.

#### Suggested Usage

Use natural language prompt and don't seperate every word with comma `,` then reinforce composition of image with additional tags.

## Prompt Examples

### Comparison Between SeaArt Infinity & Realism

**Positive Prompt**

> professional photographic shot of a beautiful woman, she has slightly wavy long ginger hair, her green eyes shining with reflection of light, she sitting on attic, sunshine comes inside from single square window, she smiling to the viewer, slightly tilted head, pastel bokeh effect, depth of field, subsurface light scattering, backlighting, dark ambient.

### Lamborghini

<figure><img src="/files/7UI6DcLh8mRveE5wPbcf" alt="" width="563"><figcaption></figcaption></figure>

**Positive Prompt**

> Cinematic photography, side-view of a Lamborghini Aventador with iridescent silver chroma in a modern luxury car studio, black studio with dim spot lighting, Big graffiti "SeaArt Studio" written on black wall with modern font, sharp rim wheels and lowered body, colorful smoke and dust particles around wheels looks totally amazing.

### This War Of Mine

<figure><img src="/files/AC1xtknYM10Pb5n8S7Is" alt="" width="563"><figcaption></figcaption></figure>

**Positive Prompt**

> Cinematic photography, eye-level view of a torn and damaged bunny plushie in wastelands, one of its eye is missing and it has bloodstain on itself, rainy and dark gothic atmosphere covered everywhere, thought-provoking war image by professional photographers.


# NOOBAI XL

NOOBAI-XL specializes in anime image generation, replicating anime characters, artists' styles, and furry content with minimal use of Lora models.

NOOBAI XL is based on the SDXL model architecture, with the Illustrious-xl-early-release-v0 as the base model. It has been trained on a large number of iterations using the complete Danbooru and e621 datasets (approximately 13,000,000 images), providing rich knowledge and excellent performance.

NOOBAI-XL is primarily used for anime image generation and is capable of reproducing the styles of thousands of anime characters and artists. It recognizes a wide range of unique concepts from the anime world and has extensive knowledge of furry content. It excels at generating tasks with minimal use of Lora models.

Try the Model: [NoobAI-XL (NAI-XL)](https://www.seaart.ai/models/detail/0a6fd94c982d214506da9355b1a5c74b)

SeaArt Exclusive Model: [WAI-Pluralistic-Noob](https://www.seaart.ai/models/detail/ct0okkde878c739ictog)

Based on the NOOB model, it has a stable basic art style compared to NOOBCollaborate with NOOB artists to explore various composition angles and colors.

## Model Usage

* It is recommended to add aesthetic and quality tags before the prompt words.\
  **Aesthetic tags:** very awa, worst aesthetic\
  **Quality tags:** masterpiece > best quality > high quality / good quality > normal quality > low quality / bad quality > worst quality
* The trigger word for a character is its character name; the trigger word for an artist's style is the artist's name
* NoobAI-XL supports special tags such as quality, aesthetics, creation year, creation period, and safety level for auxiliary use.

## Prompt

### Prompt Template

**Prompt:** very awa, masterpiece, best quality, newest, highres, absurdres&#x20;

**Negative Prompt:** NSFW, low quality, worst quality, normal quality, text, signature, jpeg artifacts, bad anatomy, old, early, copyright name, watermark, artist name, signature

### Prompt Guidelines

* It is recommended to use phrases, separated by <mark style="background-color:yellow;">", "</mark>. For example: 1girl, solo, blue hair.
* Prompts should <mark style="background-color:yellow;">not</mark> contain any underscores <mark style="background-color:yellow;">"\_"</mark>.
* If necessary, escape parentheses by adding a backslash <mark style="background-color:yellow;">"\\"</mark> before the parentheses. For example, 1girl, ganyu (genshin impact) should be written as 1girl, ganyu genshinimpactgenshin impactgenshinimpact.
* It is recommended to write prompts in logical order,&#x20;

<mark style="background-color:yellow;">for example:</mark>\
<1girl/1boy/1other/female/male/...>, \<character>, \<series>, \<artist(s)>, \<general tags>, \<other tags>, \<quality tags>\
where \<quality tags> can be placed at the front.

### Characters and Artists

* For artist styles, simply include the artist's name in the prompt without any prefixes, suffixes, or additional modifiers. It should not be "by xxx", "artist: xxx", <mark style="background-color:yellow;">just "xxx".</mark>
* For characters, use the format "character name + series". In addition to the character name, you should also add a series tag right after the character's trigger word to indicate which work the character is from.\
  For example, for the character "Ganyu" from "Genshin Impact", the prompt should be "ganyu (genshin impact), genshin impact".

## Generation Parameters

**Sampler:** Euler/Euler a\
**Steps:** 28-40\
**CFG Scale:** 3.5-5.5\
**Resolution:**

\
Resolution (Width x Height): 768x1344, 832x1216, 896x1152, 1024x1024, 1152x896, 1216x832, 1344x768\
Aspect Ratio: 9:16, 2:3, 3:4, 1:1, 4:3, 3:2, 16:9\
Notes:\
No need to set CLIP skip\
No need to use any other VAE models

<table><thead><tr><th width="133"></th><th width="106"></th><th width="103"></th><th></th><th width="109"></th><th></th><th></th><th></th></tr></thead><tbody><tr><td>Resolution (Width x Height)</td><td>768x1344</td><td>832x1216</td><td>896x1152</td><td>1024x1024</td><td>1152x896</td><td>1216x832</td><td>1344x768</td></tr><tr><td>Aspect Ratio</td><td>9:16</td><td>2:3</td><td>3:4</td><td>1:1</td><td>4:3</td><td>3:2</td><td>16:9</td></tr></tbody></table>

<mark style="color:red;">**Notes:**</mark>

No need to set CLIP skip&#x20;

No need to use any other VAE models

## Result Example

**Lora:** [AI styles dump (Illustrious/pony/noob)](https://www.seaart.ai/models/detail/667f5a285bb92b1162a13f2c245f2745)

**Prompt**

> masterpiece, best quality, good quality, nyalia,1girl, animal hands, animal ears, plana (blue archive), heart, halo, pov hands, black eyes, paw gloves, cat ears, pov, gloves, black hairband, long hair, cheek squash, hand on another's face, red halo, cat paws, hand on another's cheek, black choker, fur trim, hairband, symbol-shaped pupils, one eye closed, colored inner hair, braid, hair over one eye, multicolored hair, choker, looking at viewer, single braid, solo focus, pink hair, white background, white hair, fake animal ears, simple background, heart-shaped pupils, blush, red pupils, 1boy, upper body, closed mouth, bare shoulders, pink pupils, pink halo, tail, cheek press, cat tail, off shoulder, black gloves, 1other, two-tone hair, alternate costume, smile, animal ear fluff, sensei (blue archive), black coat, grey hair, long sleeves

**Negative Prompt**

> worst quality, bad quality, mammal, anthro, scalie, lowres, jpeg artifacts, bad anatomy, bad hands, multiple views, signature, watermark, censored, sketch, flat color, pale color, low contrast, muted color, limited palette, greyscale, monochrome, depth of field, sepia, yellow background

**Prompt**

> artist:gomzi,artist:wanke,artist:diyokama,frieren,himmel\_(sousou\_no\_frieren),sousou\_no\_frieren BREAK,rei\_(sanbonzakura),\[\[artist:fuzichoco]],year 2023,((((alternate\_costume)))),solo,1boy,1girl,couple,(lifting person:1.331),(carrying person:1.1),face to face,eye contact,arms around neck,flora background,cowbo shot,

**Negative Prompt**

> lowres,(bad:1.1),error,fewer,extra,mismatched pupils,missing,worst quality,jpeg artifacts,bad quality,watermark,weibo logo,weibo watermark,unfinished,displeasing,chromatic aberration,signature,extra digits,artistic error,username,scan,(abstract:0.909091),artist name,bad anatomy,bad hands,extra fingers,bad feet,wrong foot,wrong hand,bad leg,bad feetfingers,bad proportions,too many hands,too many fingers,text,missing fingers,extra digit,fewer digits,cropped,worst quality,low quality,normal quality,(animal, dool, chibi:1.61051),demon horns,broken horns,single horn,demon wings,extra horns,red skin,tentacle hair

**Prompt**

> 1girl,,camellya (wuthering waves),,::, (by kana616:0.8), \[(by hen-tie:1.3)|(by kagami\_(galgamesion):1.2)], (by kinokohime:1.2), (by yatsuha\_(hachiyoh):0.8), year 2024,,parted lips, looking at animal, upper body, profile, holding, centauroid, holding bouquet,,streaked hair, hair flower, multicolored hair, breasts, hair ornament, hair between eyes, jewelry, twintails, white hair,,animal ears, deer girl, flower wreath, white bird, off-shoulder shirt, shirt, sleeves past wrists, deer ears, off shoulder, head wreath, long sleeves,,blurry background, depth of field, bokeh, planted sword, red background, planted, red theme, bird, grass, flower, very awa, masterpiece, best quality, newest, highres, absurdres

**Negative Prompt**

> (worst quality, low quality, ugly:1.4), furry, (multiple\_views:2), multiple\_views, (comic, 4koma, censored, bar\_censor, mosaic\_censoring), poorly drawn hands,poorly drawn feet,poorly drawn face,out of frame,mutation,mutated,extra limbs,extra legs,extra arms,disfigured,deformed,cross-eye,blurry,(bad art, bad anatomy:1.4),blurred,text,watermark,negative\_hand-neg,((( three fingers, four fingers, six fingers, seven fingers, extra fingers, missing fingers, fused fingers, deformed fingers, ugly fingers,deformed hands, bad hands, worst time, worst hands, wrong hands, twisted hands, ugly hands,deformed toes, fused toes, missing toes,wrong feet, deformed feet, ugly feet))),unaestheticXL\_cbp62 -neg,NEGATIVE\_HANDS,


# T-Ponynai3 V6

T-ponynai3 V6 is a model tailored for generating distinct anime art with vivid colors, utilizing built-in features to simplify usage and improve style accuracy.

T-ponynai3 is an high quality pony based model. It specializes at generating anime art with certain style. Uses negative prompts and works differently from other pony based models. Model already has included vae, so you don't have to use "sdxl\_vae" as other models. Term "anime" is used on training model, which means using "anime" in your prompt will reduce style confusion and model will focus on certain anime style instead of western style cartoons. Style of model is different from general anime models and looks pretty with vivid colour selection

**Try the Model:** [T-Ponynai3 V6](https://www.seaart.ai/models/detail/34ec36d43cf42880bf985f83fe3d7d85)

## What This Model Excels At

* Crafting stylized anime arts.
* Has Pony Diffusion based special tags for better experience. Only `source_anime` and `rating tags` are recommended because model focuses on anime arts.
* Generating human/humanoid, however model still can generate pony/furry images but it's not main purpose of model. (Not recommended to use for furry/pony images.)
* Generating female/male/transgender subjects.
* Crafting SFW/NSFW content imagery.
* Crafting images with wide-range themes. (fantasy, sci-fi, medieval etc.)
* Recognizes many of Danbooru based tags, but model training is not Danbooru based.
* Recognizes only artist database of Pony Diffusion, no artist added from nai3.
* Knowledge of many well-known anime/cartoon/game characters. (Especially Genshin Impact characters)
* Knowledge of various face/body/characteristic/outfit terms.
* Can create amazing background/landscapes. (Especially V6)
* More stable than other PonyV6 based models.
* Better at hand, feet renders.
* Great lighting/shadow usage.

## Cautions For Optimal Usage

* Model uses negative prompts, but best usage is keeping it short and using recommended keywords.
* Keyword `anime` can be used to increase style of image.
* Does not require quality modifiers like `masterpiece, best quality, hd` etc.
* Make sure to use Clip Skip at `2`.
* Perfectly supports LoRAs trained with PonyV6 as the base model, yet AnimagineV3 and SDXL1.0 LoRAs can be used sometimes but not recommended.
* Do not use `sdxl_vae`, model already has included vae.
* Can't create photorealistic images but pretty good at rendering realistic cartoon images.
* Use `anime` (for nai3) and `source_anime` (for PonyV6) tag to achieve desired anime style.
* `score` tags must be used front of prompt.
* 1024px resolution is recommended, although model can generally work with most of supported SDXL resolutions.

## Recommended Settings

* Steps: 25-30
* CFG Scale: 7
* Sampler: Euler a
* Clip Skip: 2
* VAE: None, already includes vae

## Suggested Resources

### LoRA

* LoRAs based on Pony Diffusion V6 XL
* SDXL style LoRAs (can cause error/crash)

### Prompt Style

* Requires `score_9, score_8_up, score_7_up, score_6_up, score_5_up, score_4_up` or `score_9, score_8_up, score_7_up` in front of positive prompt.

<mark style="color:red;">**Note:**</mark> Full string should be more effective, try and select one you prefer.

* Model uses negative prompts, unlike the other PonyV6 based models. `(score_4, score_3, score_2, score_1)` Model uses lower score tags at negative while many Pony models don't have them.
* Use `ugly, bad anatomy, bad hands, bad feet` like terms in negative to improve accuracy of generation.
* Keyword `anime` and `source_anime` can be used to acquire desired style.
* Data selection tags can be used to make AI focus on certain contents.
* Can work with both of short and long prompts.
* Prefers tag like prompts. Model isn't Danbooru based but can understand Danbooru tags.

### Recommended Prompt Order

1. `score_9, score_8_up, score_7_up, score_6_up, score_5_up, score_4_up`
2. `source_anime, anime` & Data Selection Tags (User Preference)
3. Subject
4. Describe your suggested composition in general order. Use tag based prompts for better understanding.

### Special Data Selection Tags

* Tags can be used in both positive and negative prompts to focus on a specific data of the Model's training database.

#### Source

* Model excels in anime style, be sure to use `anime` (for nai3) and `source_anime` (for PonyV6) in your prompt. You can use other source tags too.
* source\_cartoon (Not Recommended)
* source\_pony (Not Recommended)
* source\_furry (Not Recommended)

#### Content Safety Rating

* `rating_safe`
* `rating_questionable`
* `rating_explicit`

## Prompt Examples

### Raiden Shogun

**Positive Prompt**

> score\_9, score\_8\_up, score\_7\_up, score\_6\_up, score\_5\_up, score\_4\_up, source\_anime, (anime), (front view), solo focus, 1girl, raiden shogun, \[disdain|haunted] face, mole under eye, holding, holding weapon, holding sword, (sword between breasts), (human scabbard:1.2), electricity, musou isshin (genshin impact), sakura forest, (purple sky:1.2), \[\[lightning]], aesthetic pose, ready-to-fight

**Negative Prompt**

> (score\_4, score\_3, score\_2, score\_1), ugly, bad anatomy, bad hands, (bad composition, wrong composition)

**LoRA**

* [Booba Sword (Human Scabbard)](https://www.seaart.ai/models/detail/814e556e496c9d22128d5e398d2f9ead)&#x20;

### Girl's Abyss

<figure><img src="/files/BbR28p19768t4MoHYwR7" alt="" width="369"><figcaption></figcaption></figure>

**Positive Prompt**

> score\_9, score\_8\_up, score\_7\_up, score\_6\_up, score\_5\_up, score\_4\_up, source\_anime, solo focus, sfw, ((portrait of a girl)), upper body, abstract art, surreal, touching a butterfly, floating hair, blue butterfly, glowing scenery, (falling petals, rose petals:1.2), blue theme, crimson theme, (aquamarine theme:1.2), (dark contrast:1.3), color chiaroscuro, (abyss, mysterious fields:1.3), amazing background, subsurface scattering, whimsical ethereal art

**Negative Prompt**

> (score\_4, score\_3, score\_2, score\_1), ugly, bad anatomy, bad hands, nsfw


# Counterfeit V3.0

Counterfeit is a high-quality anime model tailored for creating stylized images from close-ups to distant shots.

Counterfeit is a female focused high-quality anime model that can create stylized images from close-up portraits to long-distance shots. The model has a user-friendly prompt language that supports Danbooru tags, but "EasyNegative" embedding is essential for high-quality outputs. Additionally, it can work at different resolutions compared to other models.

<figure><img src="/files/LzpJ85T97hZ05fiv1Cnm" alt=""><figcaption></figcaption></figure>

**Try the Model:** [**Counterfeit-V3.0**](https://www.seaart.ai/models/detail/038254337d59ef522fdb64268bc28e47)

## What This Model Excels At

* Generating female/male images.
* Generating SFW/NSFW content.
* Crafting highly creative compositions.
* Crafting stylized anime style arts.
* Danbooru tags can be used while prompting, natural language prompts can be effective as well.
* Works well with various resolutions.
* Generating abstract/landscape illustrations.

## Cautions For Optimal Usage

* Highly towards to generate feminine/female people.
* Using "EasyNegative v1 or v2" embedding can increase quality of image to highest levels, highly recommended to use.
* There is no clear difference between embedding versions, "v2" is trained on "Counterfeit v3" model. User choose one of them to use, model doesn't require any other negative in most of the situations.
* There is higher possibility to get anatomical errors because model prioritize the freedom of composition. (Very bad hand generations)
* Can't recognize most of well-known characters, only generate with some of Danbooru character tags. e.g. `mona \(genshin impact\)`

## Recommended Settings

* Steps: 20-25
* CFG Scale: 8-10
* Sampler: DPM++ 2M Karras, DPM++ SDE Karras
* Clip Skip: 2
* VAE: None, kl-f8-anime2
* Upscale: R-ESRGAN 4x+ Anime6B
* Upscale Sampling Times: 10-15
* Redraw Noise Intensity: 0.6

## Suggested Resources

### Embedding

* EasyNegative
* \[TI] EasyNegativeV2 \[Textual Inversion Embedding]

## Prompt Style

* Can work with both of tag based and natural language prompts.
* "EasyNegative" or "EasyNegativeV2" trigger words must be used on negative prompt, if user going to use embeddings. Don't use them together and pick one version to use. They don't have any significant difference but this is a recommendation for embedding's training:

EasyNegative: Counterfeit v2/v2.5&#x20;

EasyNegativeV2: Counterfeit v3

* Works well with quality modifier prompts like `masterpiece, best quality, highly detailed, absurdres` etc.
* Great usage of `depth of field`
* Recognizes most of Danbooru tags.
* Does not need any other negative prompt with embedding, but they still can be used.
* `girl` is a trigger word for model. This means model will generate female focused images, if you use `girl` anywhere in your prompt.

### Recommended Prompt Order

1. Quality/Detail Terms
2. Angle/View
3. Subject
4. Subject Details
5. Additional Terms/Tags
6. Lighting/Shadow

## Prompt Examples

### Office Senior

**Positive Prompt**

> (masterpiece, best quality), highly detailed, photorealistic artwork, sfw, (from above), 1girl, brown eyes, glasses, blush, green hair, hair ornament, big breasts, office lady, formal, standing, looking up, looking at viewer, indoors, office background, office desk, office chair, (depth of field), (bokeh:1.4), (pastel colors:0.8), \[film grain], (by ligne claire:1.2), absurdres, dramatic studio lighting, soft shadows

**Negative Prompt**

> EasyNegativeV2, nsfw, child, childish

### Date At Night

**Positive Prompt**

> (masterpiece, best quality:1.3), sfw, solo, dutch angle, 1girl, red eyes, serious, frown, black hair, long hair, straight hair, crop top, midriff, cleavage, leather jacket, black jacket, black pants, outdoors, street background, (lamppost:1.2), (night), sharp focus, depth of field, (subsurface light scattering), backlighting, absurdres

**Negative Prompt**

> EasyNegativeV2, nsfw, child, childish


# Temporal Paradox Mix

It works well with both detailed and short prompts, and it performs above average when creating hands, even accurately rendering fingernails.

## **Model:** [**Temporal Paradox Mix**](https://www.seaart.ai/models/detail/f1af3436fdc289cccba0bef1b9fdaeb9)

**Creator:** Chronos dragon&#x20;

**Type:** Checkpoint&#x20;

**Introduction:** Temporal Paradox Mix is a model merge of at least 10 other merge models. It's versatile able to both female and male characters, landscapes, mechanical things like mechs, and even animals. It works with detailed longer prompts and still gets good results with short prompts too. This mix is also above average when it comes to making hands, even including finger nails

**Steps:** I use 35-40 steps personally but it should still work well at 20&#x20;

**Sampler:** Euler a, or DPM++ 2M Karras&#x20;

**Clip Skip:** 1 or 2&#x20;

**Suggested Negative embeddings:** verybadimagenegative\_v1.3, ng\_deepnegative\_v1\_75t, bad-hands-5, bad\_prompt\_version2, EasyNegative


# High-Quality AI Apps Recommendation

Discover amazing AI art tools! Transform anime to real, explore futuristic 3D, and more with these top recommendations.

1. [**anime to real**](https://www.seaart.ai/workFlowAppDetail/col9tite878c7397l90g)

<figure><img src="/files/4fWnZgmZvNE3EqquLp9H" alt="Anime to Real" width="375"><figcaption></figcaption></figure>

Author: [NekoCat](https://www.seaart.ai/user/0f3c4da03e6715a3f6ae9f8e17c43be2)

Upload a picture and click to transform anime girl into real girl. Closeup picture is prefer.

2. [**futuristic 3D style**](https://www.seaart.ai/workFlowAppDetail/cnttchte878c73cfoaj0)

<figure><img src="/files/ueRwrKA09PH1Y4W9UflL" alt="Futuristic 3D Style"><figcaption></figcaption></figure>

Author: [煮刀论AI](https://www.seaart.ai/user/260998846)

Generate a futuristic 3D book design, infused with modern technology and creative elements.

3. [**French Art Nouveau x Steampunk Art**](https://www.seaart.ai/workFlowAppDetail/cogcmgte878c73c1k9m0)

Author: [圣诞又至](https://www.seaart.ai/user/4a8bd4a44dadca93c34893f6eee0bb5d)\
Blend the elegant and romantic Art Nouveau style with the retro-futuristic aesthetics of Vaporwave to create a visual experience that is both nostalgically charming and forward-looking.

Ckpt recommendation: SDXL 吴哈柏 二次元；MeinaMix；CG Style CG 风格大模型 【Final】;&#x20;

Animagine XL V3.1


# High-Quality Character Recommendation

Recommend high-quality, intelligent characters with outstanding expressiveness and high interactivity, providing you with a lifelike and personalized experience.


# High-Quality Workflow Recommendation

Recommend an efficient and high-quality ComfyUI workflow, with a clear structure and powerful features, providing you with a smooth and professional creative experience.


# 7-FAQ

Learn about SeaArt AI, its features, copyright policies, and how to avoid generating NSFW content in this helpful FAQ.

> **Q1. What is SeaArt AI?**&#x20;
>
> SeaArt AI is a highly efficient and easy-to-use AI tool that includes multiple functions such as AI painting, AI Canvas, and ComfyUI. Users can easily generate a large number of high-quality images suitable for various scenarios without the need for professional skills. With a rich library of models and professional-level settings, combined with an intelligent recommendation system and community interaction sharing features, high-quality creation is within reach. It offers a variety of AI tools, including face swapping, AI filters, sketch-to-image, background removal, and animation generation, allowing you to quickly produce realistic, high-quality works that meet personalized needs. Whether you are a beginner or a professional, SeaArt provides a drawing method and unique style that fits your needs.

> **Q2. Can I use images I created in SeaArt, or use works from other SeaArt users for commercial purposes?**&#x20;
>
> The intellectual property rights of the content generated by you on SeaArt belong to you, and we do not prohibit you from using your own or others' works for commercial purposes. However, please note that when using others' works for commercial purposes, you need to obtain authorization from the owners. We are not responsible for any risks associated with commercial use, and you need to assume the relevant risks and responsibilities on your own.

> **Q3. What is NSFW (blurred/pink images)? How can I avoid having images tagged with this label?**&#x20;
>
> NSFW stands for "Not Safe For Work," which typically refers to content that is inappropriate for viewing in a professional or public setting. This label is often applied to emails, videos, blog posts, and forum threads that contain explicit sexual content, graphic violence, or other extreme material to prevent inappropriate exposure. To avoid having your images tagged as NSFW, you can enter NSFW in the negative prompts when generating images. This can help reduce the likelihood of creating NSFW content.

## **FAQ & Troubleshooting**

### **1. Generator Errors**

#### **1-1 What Does "Canceled by System" Mean? How to Fix It?**

● This error is usually related to specific model versions or server resources. Try using recommended models.

● Lowering generation parameters (e.g., resolution, steps).

● Avoiding peak hours.

● Contact customer service if this occurs frequently.

#### **1-2 Other Common Error Messages and Solutions**

● "Out of memory": Lower resolution, or close other programs.

● "Network error": Check network connection, or refresh the page.

#### **1-3 Common Causes of Image Generation Failure**

● Parameters set too high, device lacks resources.

● Prompts are too complex or not supported by the model.

● Network unstable.

● It's recommended to troubleshoot step by step and simplify parameters and prompts.

### **2. UI and Performance Issues**

#### 2-1 The Site Runs Slowly on My Device (Low-End Device Optimization Tips)

● Use basic models and lower generation parameters.

● Close other programs that are using resources.

● Use a PC or high-performance device.

● On mobile, use the latest version of your browser.

#### **2-2 UI Display or Operation Issues on Mobile**

● Try refreshing the page or using another browser.

● Check the network environment.

● We are working to improve mobile adaptation. Please report any serious problems.

#### **2-3 Possible Reasons for Slow Image Loading**

● Insufficient network bandwidth.

● Server is busy.

● Try switching networks or retrying later.

### **3. Account-Related Issues**

#### **3-1 What If I Get an “Invalid Email” or Account Expired Message When Logging In?**

● Double-check that your email is entered correctly.

● If the email is correct but the login still fails, it may be due to account issues or mistaken suspension. Please contact customer service.

● If your account suddenly becomes inaccessible and data is lost, please provide relevant details, and we will assist with account recovery.

#### **3-2 What If My Account Suddenly Loses History or Credits?**

● First, check if you changed your login method or email.

● If the data is indeed lost, contact customer service promptly and provide your account info and relevant screenshots. We will help investigate and recover the data.

#### **3-3 How to Deal with Frequent Re-login Requests?**

● Clear your browser's cache and cookies.

● Try a different browser or device.

● Check if your account is being logged in from multiple locations.

● Contact customer service if the problem persists.

### **4. Content Moderation Issues**

#### **4-1. What Is the Site's Content Moderation Policy?**

● It is strictly prohibited to generate content that is illegal, violent, pornographic, or infringing.

● NSFW content is strictly filtered.

#### **4-2. My Work Is Mistakenly Flagged as NSFW. How Do I Appeal?**

● Contact customer service if this occurs frequently.

#### **4-3. Why Is Similar Content Sometimes Approved and Sometimes Not?**

● As the moderation system combines AI and human review, misjudgment may occur.

● We are continually optimizing the moderation algorithm and welcome users to report specific cases to help improve the system.

### **Contact & Feedback**

#### **1. How to Get Help on Discord**

● Join the official Discord server, navigate to the #Feedback & Suggestions channel, and describe your issue.

● A dedicated support staff will assist you directly.

#### **2. How to Report Bugs or Contact Customer Service**

● Join our official Discord server, navigate to the #bug-reports channel, describe your issue in detail, and include screenshots whenever possible.


# SeaArt AI 使用ガイド

SeaArt AIの使い方ガイドをご利用ください。

## SeaArtへようこそ！

SeaArtは、無料で利用できる最先端のAIアートジェネレーターです。活気に満ちたAIコンテンツコミュニティに参加し、1,000,000以上のモデルやスタイルを探索しましょう。アートからイラスト、絵画まで、SeaArtはあなたの創造力を簡単に引き出します。

今日から無料で始めて、あなたのクリエイティブなワークフローを向上させましょう！

**SeaArt APP:** [*https://seaart-all.onelink.me/qoIB*](https://seaart-all.onelink.me/qoIB/DC)

**Discord:** [*https://discord.com/invite/gUHDU644vU*](https://discord.com/invite/gUHDU644vU)

> **ご質問がある場合は、Discordサーバーを通じてお問い合わせください。**

## <mark style="background-color:yellow;">公式SNSのフォローをお願いします：</mark>

* Instagram: [@seaartai](https://www.instagram.com/seaartai/)
* X: [@SeaArt\_Ai](https://x.com/SeaArt_Ai)
* TikTok: [@seaart.ai](https://www.tiktok.com/@seaart.ai?lang=en)
* Youtube: [@SeaArt](https://www.youtube.com/channel/UC-hWsHj8797Sv2sf20-6-3g)
* Facebook:[ @SeaArt\_Ai](https://www.facebook.com/SeaArtAiOfficial)
* Reddit: [@SeaArt\_Ai](https://www.reddit.com/r/SeaArt_Ai/)

## **🐾**[SeaArt.AI](https://www.seaart.ai/ja)の始め方！！！

## [SeaArt.AI](https://www.seaart.ai/ja)とは？

[SeaArt.AI](https://www.seaart.ai/ja)は、誰でも利用できる無料のオンラインAIアート作成プラットフォームです。インターネットに接続していれば、どこからでも[SeaArt.AI](https://www.seaart.ai/ja)のサービスをご利用いただけます。PC、Androidデバイス、Appleデバイスなど、どのデバイスからでも、快適で魅力的なアート制作体験をお楽しみいただけます。\ <br>


# 1-基本ページ

SeaArt AIの基本を学び、今すぐAIアート制作を楽しみましょう。

## 新規ユーザー登録

**ステップ1：**&#x30C8;ップメニューバーの右側にある「新規登録」をクリックし、ログイン画面を開きます。

<figure><img src="/files/jWHlgikQc0CNPAFTmUZu" alt=""><figcaption></figcaption></figure>

**ステップ2：** 適切な方法を選んでログインまたは登録してください。

<figure><img src="/files/hkrkpwAIe8JDuG5BYkx5" alt=""><figcaption></figcaption></figure>

## 登録／ログインに関するよくある質問

**1.メール認証のよくある質問（認証メールが届かない場合）**

● メールアドレスが正しく入力されているか確認してください

● 迷惑メールフォルダや広告メールフォルダを確認してください

● 数分待ってから、再度メールボックスを更新してください

● それでも届かない場合は、よく使うメールアドレス（Gmail、Outlookなど）に変更してみてください

● 何度試しても届かない場合は、カスタマーサポートにお問い合わせください

**2.サードパーティアカウントでのログイン方法（Googleなど）**

ログインページで「Google」ボタンを選択し、案内に従って認証を完了してください。ログインに失敗した場合は、ブラウザのキャッシュをクリアする、ブラウザを変更する、またはネットワーク環境を切り替えてみてください。それでもログインできない場合は、カスタマーサポートに具体的なエラーメッセージを添えてご連絡ください。

**3.サードパーティログインが失敗する場合の基本的なトラブルシューティング**

● ブラウザのキャッシュとCookieをクリアする

● ブラウザを変更するか、シークレットモード（プライベートブラウジング）を使用する

● Googleアカウントの状態が正常か確認する

● それでも失敗する場合は、エラー画面のスクリーンショットを撮ってカスタマーサポートにご連絡ください

**4.パスワードを忘れた場合／パスワード変更**

ログインページで「パスワードを忘れた場合」をクリックし、登録時のメールアドレスを入力してください。システムからパスワードリセット用のメールが送信されます。メールの案内に従って新しいパスワードを設定してください。

<figure><img src="/files/dFyPzjr1MFb7s6k3lyGp" alt=""><figcaption></figcaption></figure>

## 初心者向けクイックスタート

**最初の作品を作成する手順**

「作成」ボタンをクリック → 画像／動画を選択 → モデルを選択 → プロンプトを入力 → 「創作」ボタンをクリックして終了です。

## モデルの選択と使用

**1.モデルの閲覧・検索方法**

作成フロー内のモデル選択欄では、キーワード検索やスタイル、タイプなどのフィルターを使ってモデルを閲覧できます。目的のモデルが見つからない場合は、ページを再読み込みするか、そのモデルが公開停止になっていないかを確認してください。

**2.モデルスタイルの紹介とサンプル**

　◈ SDXL：リアルな描写や写実的なスタイルに最適です。

　◈ Animeシリーズ：アニメや二次元スタイルの作品に適しています。

　◈ Illustriousシリーズ：イラストやアート系の表現に向いています。

　◈各カテゴリで好みのスタイルを選び、モデルページの「おすすめ」で生成例を確認できます。

**3.目的に合わせたモデルの選び方**

生成したい作品のイメージに合わせて、適したスタイルのモデルを選びましょう。

モデルの説明文、コミュニティのおすすめ、サンプル画像などを参考にするのがおすすめです。思ったような結果が得られない場合は、モデルを変更したり、プロンプトを調整してみてください。

ホームページ内の創造コンテストでは、新着モデルや人気モデルの使い方・レビューが随時更新されます。チェックしておくと便利です。

**4.モデルが読み込めない／検索できない場合の主な原因と対処法**

◈ ネットワークの不安定やサーバーの混雑が原因の場合があります。ページを更新するか、時間をおいて再試行してください。

◈ モデルが非公開またはメンテナンス中の可能性があります。おすすめモデルの使用を検討しましょう。

◈ 何度試しても読み込めない場合は、カスタマーサポートまでお問い合わせください。

## ページの紹介

### **ホームページの構成**

左側のナビゲーションバーから、さまざまな機能を探索できます。

ウィンドウをスクロールすると、期間限定イベントや新機能のリリース情報を確認できます。

ウィンドウをスクロールすると、多機能AI創作コミュニティで、様々なクリエイティブな使い方を確認できます。また、【おすすめチャット】で紹介されている人気のAIキャラクターと次元を超えた会話を楽しんだり、キャラクターセンターで自分専用のAIキャラクターを作成したりすることもできます。

<figure><img src="/files/nNj3kTXacvAUYcb29iyF" alt=""><figcaption></figcaption></figure>

さらに下にスクロールすると、他のユーザーが投稿した作品集もご覧いただけます。

#### AIアート

> アクセス方法：右上の「作成」ボタンの横にある「ｖ」アイコンにマウスを合わせると、画像生成、動画生成、AIキャラクター、音声、ワークフローなどの機能が表示されます。お好みの機能を選択してお試しください。

<figure><img src="/files/C1yg6FnFcZUAc5yQoVTM" alt=""><figcaption></figcaption></figure>

#### AIクリエイティブツール

選択したモデルによって、画像生成では主にAI画像アレンジ、Img2Img（画像から画像生成）、参考画像などのAIクリエイティブツールが利用できます。

<figure><img src="/files/HI0IundVl51fKTop7WmZ" alt=""><figcaption></figcaption></figure>

#### AIアプリ

* **動画作成：** 静止画を動画に変換します。

<figure><img src="/files/mmiIMUsF6wkEezy6ZG0E" alt="A bird on a branch" width="240"><figcaption></figcaption></figure>

* **AI画像拡張：**&#x753B;像をアップロードした後、拡張したい部分をドラッグで選択します。元の画像と同じスタイルのモデルを選択し、拡張したい部分に関連するプロンプトを入力するだけです。
* **AIフェイススワップ：** AI顔入れ替えとも呼ばれます。元の画像をアップロードした後、右側に置き換えたい顔を追加します。
* **アップスケール：** 画像の詳細を変更せずに解像度を上げます。
* **人物修復：** 画像内の顔、手、体を修復するために使用します。元の画像と同じスタイルのモデルを選択することをお勧めします。

<figure><img src="/files/U0y56ZiTniS3QrEGYgvR" alt="Before and after comparison of image repairing"><figcaption></figcaption></figure>

* **画像キーワード摘出：**&#x753B;像をアップロードすると、プロンプトが表示されます。
* **前処理プレビュー：** 生成済みの画像をアップロードし、コントロールネットのタイプを選択して「元の画像」を逆算します。
* **プロンプト最適化：**&#x4F5C;成欄にプロンプトを入力してこの機能をクリックすると、AIが自動的に内容を最適化してくれます。

<figure><img src="/files/QXnM9ngjvrR37Xs1y9Hx" alt=""><figcaption></figcaption></figure>

### ホームページ

**アクセス方法：**&#x5DE6;上の<mark style="background-color:yellow;">「</mark>[ホームページ](https://www.seaart.ai/ja)<mark style="background-color:yellow;">」</mark>をクリックしてください。

◈ **検索バー：**&#x4F5C;品／モデル／AIキャラクター／AIアプリ／投稿などを検索することができます。

<figure><img src="/files/qOYQLzhhIKRoUXR7BmRH" alt=""><figcaption></figcaption></figure>

◈ **オンラインサポート：**&#x753B;面右下のサポートボタンをクリックすると、オンラインカスタマーサービスに連絡できます。対応までにお時間をいただく場合がありますので、しばらくお待ちください。

### 作品詳細ページ

**開き方：**&#x8A72;当する作品をクリックします。

詳細ページでは、画像生成に使用されたプロンプト、モデルやその他の詳細情報などを確認できます。

また、画像の高解像度化、バリエーション（V）生成、背景削除などの機能にも対応しています。

バリエーション（V）：Img2Imgページに移動し、画像の細部を編集できます。

作品を報告する場合は、このボタンをクリックします。

<figure><img src="/files/KXkdZdwn8uvlNPBMJ0NU" alt="" width="188"><figcaption></figcaption></figure>

作品を共有またはダウンロードする場合は、このボタンをクリックします。

<figure><img src="/files/DArOQXCfO2d5wsjX4NVZ" alt=""><figcaption></figcaption></figure>

### 個人ページ

<mark style="background-color:red;">アバターにマウスを合わせると、以下の設定が表示されます：</mark>

◈ **無制限モード：**&#x7121;制限モードを有効にすると、出血・暴力・性的表現など、表示に適さない画像を非表示にできます。

◈ **言語：**&#x30B5;イトの表示言語を変更できます。

<figure><img src="/files/fKoM4w7QSt1ASuYEbbpd" alt=""><figcaption></figcaption></figure>

◈ **設定：**

**個人設定：**

* 名前変更
* ブックマークサイトのタグ設定

<figure><img src="/files/S4kaFSDyvqcLMJtCnSBj" alt=""><figcaption></figcaption></figure>

**システム設定**

* 創作完了後、作品を自動的にコミュニティに投稿する
* 履歴は自動的に個人センターに保存される
* プロンプトをデフォルトで表示する

<figure><img src="/files/DL7erRHkUuUAHXAmBL4L" alt=""><figcaption></figcaption></figure>

<mark style="background-color:red;">右上のプロフィールアイコンをクリックして個人センターにアクセスできます。</mark>

**お気に入り：**&#x304A;気に入りに追加したコンテンツを確認できます。

**概要：** 自身の全作品をまとめて表示します。

**作品：** 自分が作成した画像や動画を確認できます。

* 目アイコン： 成人向けコンテンツの表示／非表示を切り替えられます。
* 管理ボタン：動画制作、投稿、作品の整理、ダウンロード、削除などの操作ができます。
* タイムライン： 時間順に作品を確認できます。
* フィルター： 条件別に作品を絞り込みできます。

左下に、ユーザーIDと個人招待コードが表示されます。

<figure><img src="/files/VtB6gl91v4JXZNSyL1u2" alt=""><figcaption></figcaption></figure>

**クリエイター収益センター：**&#x73FE;金報酬を獲得できます。詳しいルールはこちらをご覧ください。

{% content-ref url="/pages/MvgHDJQdi9shrivEb3iU" %}
[SeaArt.AIクリエーター奨励プログラム](/guide-1/ri-ben-yu/6-naibento/seaartaikuritpuroguramu)
{% endcontent-ref %}

### タスク／消費記録

左側の「イベントセンター」 → 「タスク」をクリックすると、タスクセンターにアクセスできます。ここでは、毎日のタスク、毎週のタスク、一回限りのタスクを確認できます。

<figure><img src="/files/Vkhmxp2KbCAHMDz8nyXK" alt=""><figcaption></figcaption></figure>

「コイン獲得・利用履歴」で、スタミナ／コインの累計収入と累計支出を確認できます。

<figure><img src="/files/5tsp6et6fAZdU17wQetz" alt=""><figcaption></figcaption></figure>

### SeaArtショップ

**アクセス方法：**&#x53F3;上のVIPアイコンをクリックすると、SeaArtショップにアクセスできます。

**SeaArtショップについて**

**スタミナ：**&#x6BCE;日自動的にリフレッシュされます。基本ユーザーは1日あたり150スタミナを獲得でき、VIPユーザーはより多くのスタミナを獲得できます。

**コイン：**&#x671F;限切れになることはありません。作成中、スタミナが最初に使用され、スタミナが不足している場合はクレジットが使用されます。

**VIP特典：**&#x56;IPスタンダードコース以上は無料で無制限に創作できます。その他の特典については、下記の画像をご覧ください。

<figure><img src="/files/6HoHWPuyB7RpXs4Iw7ce" alt=""><figcaption></figcaption></figure>

**コイン購入：** コインはモデルトレーニング、AI創作、有料モデルの使用に利用でき、有効期限はありません。

<figure><img src="/files/nCBNfv1cyEq3HQp4AeKV" alt=""><figcaption><p>具体的な価格はウェブサイトに準じます</p></figcaption></figure>

<mark style="background-color:red;">**コイン**</mark>&#x20;

コインは期限が切れたりリセットされたりせず、SeaArtショップで購入できます。また、タスクを完了したり、コミュニティやウェブサイトのイベントに参加したりすることで、コインを獲得することもできます。

**a. SeaArtショップでコインを購入**

ページ右上の「VIP」アイコンをクリックすると、SeaArtショップで価格を確認してコインを購入できます。

**b. タスクを完了してコインを獲得**

左側のナビゲーションバーから「イベントセンター」→「タスク」をクリックすると、タスクリストを確認できます。対応するタスクを完了して報酬を獲得してください。

**c. コミュニティやウェブサイトのイベントに参加**

ウェブサイトやDiscordで定期的にイベントを開催しています。コミュニティ通知 [#announcement](https://discord.com/channels/1089843669944778822/1089854034866884609)に注意してください。報酬を獲得する機会をお見逃しなく！

<mark style="background-color:red;">**スタミナ**</mark>

無料ユーザーは1日150スタミナを獲得できます。VIPユーザーはレベルに応じて、1日あたり300／700／2100／3500スタミナを無料で獲得できます。

スタミナは毎日UTC 0:00にリセットされます。右上のアバターにマウスを合わせると、残りのスタミナ数を確認できます。

スタミナは日々の創作に使用できます。ここで確認できます。

<figure><img src="/files/k6zmh8Wgn7ZKequIwriO" alt=""><figcaption></figcaption></figure>

### お問い合わせ

**アクセス方法：**&#x95A2;連アイコンをクリックして、公式コミュニティやSNSに参加できます。

スタッフがオンラインで皆様のご質問にお答えします。また、SeaArtの最新アップデート情報をいち早くお届けします。AI専門家と交流し、AI技術について意見を交換しましょう！

<figure><img src="/files/amQEIsTZDaaVt7dZ6Wqn" alt=""><figcaption></figcaption></figure>


# 2-基本機能

Explore these basic functions of SeaArt, and learn how to use them to create stunning AI visual content.


# 2-1 Text to Image

テキストから画像への生成とは何か、そして簡単なステップバイステップの指示でAIテキストから画像生成ツールを使用する方法を学びましょう。

> SeaArtを使用中に、コントロール効果が理想的でない、または描画結果が追加したプロンプトを反映していないといった問題に直面したことがありますか？この記事では、テキストから画像への操作方法を包括的に紹介し、効率的なプロンプトワードの書き方を習得するための戦略を説明します。

### Text to Imageへとは何ですか？

SeaArt AIには、<mark style="background-color:yellow;">Text to Image、</mark>[<mark style="background-color:yellow;">Image to Image</mark>](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/2-2-img2img)<mark style="background-color:yellow;">、コントロールネット</mark>の3つの描画モードがあります。テキストから画像モードには、デフォルト、SDXL、スタジオの3つの方法が含まれます。

<figure><img src="/files/eUb8mjGpzNlN1ZOzZTDa" alt=""><figcaption></figcaption></figure>

描画の基本手順は、<mark style="background-color:red;">モデルの選択→プロンプトの入力→パラメータの設定→生成です。</mark>

モデルはスタイルを決定し、プロンプトは画像の内容を定義し、パラメータは画像の設定特性を細かく調整します。

### プロンプトの詳細説明

**プロンプトとは何ですか？**

AIをより効果的に誘導するために、モデルの行動を制約するためにポジティブまたはネガティブなフィードバックを提供する方法が探求されてきました。この誘導情報はプロンプトと呼ばれ、人間とAIの橋渡しの役割を果たします。

**プロンプトの基本構文**

プロンプトには、望ましい画像内容を記述するポジティブプロンプトと、画像に望ましくない内容を示すネガティブプロンプトが含まれます。

プロンプトの編集に慣れていない場合は、<mark style="background-color:yellow;">ツール - プロンプト提示</mark>スタジオをクリックして、効率的なプロンプトの組み合わせを迅速に構築できます。

<figure><img src="/files/wzWYYbirgUqCpGZfGAni" alt=""><figcaption></figcaption></figure>

特定の詳細を理解しにくいモデルがある場合（例えば、手の構造）、ネガティブプロンプトはこれらの要素を避け、画像の品質を向上させるのに役立ちます。

例えば、以下を含めます：（bad hands, bad anatomy, bad body, bad face, bad teeth, bad arms, bad legs, deformities: 1.3）

<figure><img src="/files/UsQI3U6jWnm7AH21WxuR" alt=""><figcaption></figcaption></figure>

プロンプトの入力：**自然言語/フレーズ形式**&#x20;

**自然言語**：黒髪の少女が踊る&#x20;

**フレーズ形式：**&#x5C11;女、黒髪、踊る

プロンプトの役割は、モデルの描画プロセスを誘導し、支援することであり、厳密な要件ではありません。入力が単なるカジュアルな文であっても、モデルはあなたのために画像を作成することができ、その結果はかなり良い場合もあります。

<mark style="color:red;">\*豊富なプロンプトは、最終出力効果をより良く制御できます。後の微調整プロセスでは、特定のキーワードを迅速に変更し、その影響を確認することができます。</mark>

### ユニバーサルプロンプトフォーミュラ

効果的なプロンプトは、AIアートジェネレーターにタスクを割り当てるようなものです。指示が曖昧で、「絵をデザインする」というだけで要素や目的を指定しない場合、結果は予測不可能になることが多いです。したがって、詳細で具体的な指示は、結果の品質と関連性を大幅に向上させることができます。

例えば、プロンプトが単に「少女」と入力するだけでは、少女の服装、シーン、カメラアングルなどは言及されておらず、AIはトレーニング中のモデルの履歴に基づいてしか実行できません。モデルの能力のおかげで、得られる描画結果は依然としてかなり良いですが、画面の内容に特定の要件がある場合、その効率は非常に低いです。

他の記述的な言葉を追加すると、画像ははるかに安定します。

<figure><img src="/files/g6Ts2LpcG1NbajtPQRM4" alt=""><figcaption></figcaption></figure>

理想的なプロンプトフォーミュラには、<mark style="background-color:yellow;">主な内容、環境背景、構図、画像設定、参照スタイルなどの要素が含まれており</mark>、それぞれが描画結果に異なる程度の影響を与えます。

<mark style="color:red;">\*このフォーミュラは参考であり、すべてのプロンプト作成に対する厳格なルールではありません。まず主な内容の影響を特定し、その後、個々のニーズに応じて詳細を最適化します。</mark>

<figure><img src="/files/9szNlHMfzKIqjLVnpAWQ" alt=""><figcaption></figcaption></figure>

1. **主な内容：**&#x4E3B;な内容は主題<mark style="background-color:yellow;">（人や動物）の衣装、表情、動作、物の素材などを説明します。</mark>複数の主題を一緒に生成することは問題を引き起こす可能性があります。それぞれの主題を個別に作成し、その後ControINet生成を使用して統合することをお勧めします。
2. **環境背景：**&#x74B0;境背景は、<mark style="background-color:yellow;">空の色、周囲、照明、色調などのシーンと補助要素を設定し、</mark>画像の雰囲気を強調し、テーマを際立たせます。
3. **構図のショット：**&#x69CB;図は、カメラのアングルと視点を調整し、<mark style="background-color:yellow;">被写界深度の強調やオブジェクトの配置などで視覚的なインパクトを大幅に高めます。</mark>
4. **画像設定：**&#x753B;像設定には、<mark style="background-color:yellow;">ディテールの豊かさ、写真の品質、映画的な効果など、視覚的な表現力を高める用語が含まれます。</mark>画像の解像度とディテールレベルは主にサイズによって決まり、Upscaleなどの後処理技術でさらにディテールが強化されます。
5. **参照スタイル：**&#x671B;ましい芸術スタイルやムードを説明します。例えば、<mark style="background-color:yellow;">アーティストの名前、芸術技法、時代、色などを挙げます。</mark>しかし、画像のスタイルは主にモデルによって決定されます。モデルが特定の芸術スタイルのキーワードでトレーニングされていない場合、それらを理解できないかもしれません。特定のスタイル要件がある場合、単にプロンプトを使用するよりも、そのスタイルでトレーニングされたモデルを使用する方が良い結果が得られることがあります。

創作：プロンプトの記述が難しい場合は、ホームページのAI生成画像からインスピレーションを得て、既存のパラメータとプロンプトワードをワンクリックで再利用して、作成プロセスを簡素化できます。

### 強調プロンプト

強調プロンプトは、カッコや数値を使用して特定のプロンプトの重みを制御します。重みの値が高いほど、モデルはそのプロンプトを優先し、その部分のレンダリングに集中します。結果として、最終的な画像には対応する情報がより多く反映されます。逆に、強調が少ない場合、その内容の表現が少なくなります。

重みを増加させる方法の一つは、カッコを<mark style="background-color:yellow;">使用することです</mark>。<mark style="background-color:yellow;">もう一つは数値を直接入力することで</mark>、後者が一般的に使用されるアプローチです。

プロンプトワードの重みを制御するためのカッコには3種類あります：

* 丸カッコ ( )：各層が元の重みを1.1倍に増加させます。
* 角カッコ \[ ]：各層が元の重みを0.9倍に減少させます。
* さらに、カッコは複数の層を重ねることができ、それぞれの層が一定の係数で重みを乗じます。

<figure><img src="/files/fOKNrwEipP4tfzCCZXeQ" alt=""><figcaption></figcaption></figure>

例えば、デフォルトでは、少女の服装は黄色とオレンジの組み合わせになります。しかし、「(((オレンジのコート)))」を使用すると、カッコが強調を示し、モデルのオレンジのコートの描写が強化され、最終的な画像ではコートにオレンジがより多く反映されます。

<figure><img src="/files/36jCUOQE5IXOQUbbsZpx" alt="" width="563"><figcaption></figcaption></figure>

逆に、「\[\[オレンジのコート]]」を使用すると、角カッコが強調を減少させ、オレンジの要素が減少します。モデルは残りのキーワード「((黄色のコート))」を優先し、コートがより黄色く表示されます。

<figure><img src="/files/hbF18cU4yYPyqU8RJ1pl" alt="" width="563"><figcaption></figcaption></figure>

数値を直接入力して重みを制御します。

例えば、デフォルトでは髪の色は緑と赤で表示されます。「(green hair)」の後に重みを0.9と設定すると、緑の髪の部分の重みが元の値の0.9倍に減少することを意味します。同様に、緑の髪の重みを増やしたい場合は、1.1と入力するだけです。

<figure><img src="/files/6wGoZBhikReqsddRaOvX" alt=""><figcaption></figcaption></figure>

<mark style="color:red;">\*キーワードの重みの強調は0.1から100まで変動することができますが、極端な重みの値による影響の偏りを考慮して、最適な画像結果を得るためには重みを0.5から1.5の間に保つことをお勧めします。</mark>

<mark style="background-color:red;">具体的なパラメータ設定については、ここをクリックして詳細を確認してください。</mark>

{% content-ref url="/pages/rNAxYQXt8yBJlPmXAKfD" %}
[4-パラメーター](/guide-1/ri-ben-yu/4-paramt)
{% endcontent-ref %}


# 2-2 Img2Img

画像を変換する準備はできましたか？Img2Imgへの世界に踏み込み、そのパラメータとワークフローを学びましょう。

> AIペインティングの実際の応用では、モデルによって生成される初期画像の不確実性のため、出力画像の実際のコントロール性は高くありません。そこで、「Image to Image」機能を使用して、画像を希望する方向に修正し、生成画像のコントロール性を向上させることができます。

## Img2Imgとは何ですか？

「Image to Image」機能は、既存の画像とテキストの説明を組み合わせて新しい画像を生成するAIベースの画像生成技術です。この技術は、ユーザーの特定のニーズに応じて画像とテキストのプロンプトを混合することで、新しい視覚コンテンツを作成できるため重要です。

簡単に言えば、プロンプトワードと参照画像の情報を考慮しながら描画するプロセスが、Image to Imageを構成します。

## Img2Imgへのパラメータの分析

スマート分析

提供された画像に基づいて、<mark style="background-color:yellow;">適合するプロンプトおよび画像に合ったモデルを自動的に推測します。</mark>ただし、インテリジェント分析によって生成されたプロンプトには誤ったプロンプトが含まれている可能性があるため、手動で再度スクリーニングすることをお勧めします。この機能は主にプロンプトワードの記述の参考として役立ちます。

Img2Imgへのワークフロー

<mark style="background-color:red;">ワークフロー：参照画像のアップロード - モデルプロンプトの設定 - パラメータの設定 - 生成</mark>

* 参照画像をアップロードした後、インテリジェント分析を開き、プロンプト、モデル、画像サイズを自動的に入力します。実際のニーズに応じてプロンプトを再調整することをお勧めします。パラメータ設定は初期画像の生成と同じです。最後に画像を生成するをクリックすると、AIが参照画像とユーザーの指示に基づいて新しい画像を作成します。

<mark style="color:red;">\*再描画の範囲が大きいほど、元の画像との違いが大きくなります。通常、0.4〜0.8の間に設定します。</mark>

**ノイズ除去強度：**&#x3053;のパラメータは、元の画像に基づく再描画プロセスにおける分岐の程度を制御します。値が高いほど、モデルは再描画プロセスで自由度が高くなり、描画結果と元の参照画像の違いが大きくなります。

<figure><img src="/files/alYpkWcDWU2Zk4lNn2DX" alt=""><figcaption></figcaption></figure>

ノイズ除去強度が高すぎると、描かれた画像の内容を元の画像と関連付けることが難しくなるため、再描画の範囲の値は<mark style="color:red;">通常0.4〜0.8に保ちます。</mark>

**部分的な再描画**

部分的な再描画は、画像内の特定の領域を修正および調整することを可能にします。この機能は特にローカルなディテールを微調整するのに適しています。修正内容を誘導するためにプロンプトの追加も必要です。画像の大部分の内容に満足しているが、いくつかの詳細要素を調整する必要がある場合に使用します。

画像をアップロードした後、右側のブラシをクリックして部分的な再描画エリアに入ります。次に、画像を塗りつぶすことができます。塗りつぶした後、プロンプトボックスに塗りつぶしたエリアのプロンプトを入力します。

<figure><img src="/files/zMR9BjnIgoT7LCde1egz" alt=""><figcaption></figcaption></figure>

部分的な再描画を使用した後、選択したエリアのみが再描画され、他のエリアは変更されません。

<figure><img src="/files/BBDl34hDeawI4KwxVZ0i" alt=""><figcaption></figcaption></figure>


# 2-3 コントロールネット

ControlNetでAI画像生成をマスターしましょう。プリプロセッサ、動作原理、および美しいAIアートを作成する方法について学びます。

## ControlNetとは何ですか？

ControlNetは、AI画像生成を制御するためのプラグインです。「Conditional Generative Adversarial Networks（CGANs）」という技術を使用して画像を生成します。従来の生成的敵対ネットワークとは異なり、ControlNetでは、線画をアップロードしてAIに着色させたり、キャラクターのポーズを制御したり、画像の線画を生成するなど、生成される画像を細かく制御できます。&#x20;

<figure><img src="/files/zdJT19Zg9h3U5eB2Si5r" alt=""><figcaption></figcaption></figure>

従来の描画モデルとは異なり、完全なControlNetは<mark style="background-color:yellow;">プリプロセスモデルControlNet</mark>の2つの部分で構成されます。

プリプロセスモデル：元の画像から空間的意味情報を抽出し、線画や深度マップなどの視覚的プレビュー画像に変換します。

ControlNet: 線や被写界深度などの基本的な構造情報を処理します。

<figure><img src="/files/VXwG3XLA5scTtsJrR8iX" alt=""><figcaption></figcaption></figure>

### **Canny**

基本情報

Cannyモデルは、主に入力画像のエッジ情報を識別し、アップロードされた画像から正確な線画を抽出することができます。その後、指定されたプロンプトに基づいて、元の画像の構成に一致する新しいシーンを生成します。

<figure><img src="/files/KFSxtlRsFhHP1x4JxdaS" alt="Before and after comparison of using the Canny model to process a 3d cartoon image" width="563"><figcaption><p>原図 / プリプロセッサ</p></figcaption></figure>

プリプロセッサ:

**canny:** ハードエッジ検出。

**invert:** 線画の色を白背景の黒線に反転させます。

<figure><img src="/files/kEWN0hDjhHSvGER4onU6" alt="Before and after comparison of using the Canny model to process a sketch image" width="563"><figcaption><p>原図 / invert</p></figcaption></figure>

invertはCannyに固有のものではなく、ほとんどの線画モデルと組み合わせて使用できます。Line ArtやMLSD t認識などのControlNetタイプを選択すると、invertが使用可能です。

操作方法

操作手順：

<mark style="background-color:yellow;">画像をアップロード - モデルを選択 - ControlNetタイプを選択 - プロンプトを入力 - 生成</mark>

スマート分析：

画像のプロンプトとモデルを逆推論します。元の画像と異なるスタイルを希望する場合は、インテリジェント分析をオフにすることをお勧めします。

パラメータ設定：

<mark style="background-color:yellow;">プリプロセッサ像度</mark>

プリプロセッサ像度はプレビュー画像の出力解像度に影響します。画像のアスペクト比は固定されており、デフォルトの出力は1倍の画像です。解像度の設定は基本的にプレビュー画像の横サイズを決定します。例えば、元の画像サイズと目標画像サイズの両方が512x768の場合、プリプロセッサ像度を128、256、512、1024に設定すると、プリプロセッサ画像のサイズはそれぞれ128x192（元の0.25倍）、256x384（元の0.5倍）、512x768（元のサイズ）、1024x1536（元の2倍）に変わります。

<mark style="color:red;">一般的に、解像度の設定が高いほど、生成される画像の詳細が豊かになります。</mark>

<mark style="color:red;">\*</mark>時々、前処理検出画像と最終画像のサイズが一致しない場合、描画された画像が破損し、最終描画の人物のエッジに明確なピクセル化が生じることがあります。

<figure><img src="/files/sm9L1WTX2NbEBEg2X3Zp" alt=""><figcaption></figcaption></figure>

<mark style="background-color:yellow;">コントロールウエイト</mark>

ControlNetの強度を決定します。強度が高いほど、画像効果に対する制御が強くなり、生成された画像が元の画像に近くなります。

<figure><img src="/files/UdIW0t0O8HnCZdnwCvWH" alt="" width="506"><figcaption></figcaption></figure>

コントロールモード

ControlNetとプロンプトワード間の重みの割合を切り替えるために使用されます。デフォルト設定はバランスです。

よりプロンプトに従って画像を生成します: 制御図の効果が弱まります。

より前処理画像に従って画像を生成します: 制御図の効果が強まります。

<mark style="background-color:yellow;">生成結果</mark>

生成された結果から、基本的な構成は元の画像と全く同じですが、詳細は完全に異なることがわかります。髪の色、顔の詳細、衣服など、他の変更が必要な場合は、キーワードやパラメータを調整して望む効果を得ることができます。

<figure><img src="/files/Dxz3BDz6o4f8ArZDmzPF" alt="Four examples of AI-generated superman images with different models"><figcaption></figcaption></figure>

### **OpenPose Full**

基本情報

OpenPose Fullは、<mark style="background-color:yellow;">人間の体の動きや表情</mark>の特徴を正確に制御することができます。これは、1人のポーズを生成するだけでなく、複数人のポーズを生成することも可能です。

OpenPose Fullは、<mark style="background-color:yellow;">頭、肩、肘、膝などの人間の体の重要な構造点</mark>を識別し、服装、髪型、背景の詳細を無視しながら、ポーズや表情を忠実に再現します。

<figure><img src="/files/MYwiP9qMEqM7O4SyBnNx" alt=""><figcaption></figcaption></figure>

**プリプロセッサ**

人間のポーズ認識

デフォルトのプロセッサはopenposeシリーズから来ており、<mark style="background-color:yellow;">openpose、face、faceonly、full、hand</mark>があります。これらの5つの前処理装置は、それぞれ顔の特徴、手足、手、その他の人体構造を検出するために使用されます。

<figure><img src="/files/kPhHjERybWqsYAlUnx19" alt=""><figcaption></figcaption></figure>

動物のポーズ認識

動物\_openposeプロセッサの使用をお勧めします。これは、control\_sd15\_animal\_openpose\_fp16などの専門的な前処理モデルと組み合わせて使用できます。

<figure><img src="/files/A8Y3EgtoGoeq9j9koHfP" alt="" width="563"><figcaption></figcaption></figure>

一般的には、デフォルトのopenpose\_full前処理装置を使用するだけで十分です。

### **Line Art**

基本情報

線画も画像からエッジの線画を抽出することに関するものですが、その使用ケースはより具体的で、リアルとアニメの2つの方向があります。

<figure><img src="/files/1H6lGDeXzP1ThbmIGi6i" alt=""><figcaption></figcaption></figure>

**プリプロセッサ**

Line Art

リアルな画像により適しており、抽出された線画はより復元的で、検出中にエッジの詳細をより多く保持し、そのためコントロール効果がより顕著です。

<figure><img src="/files/USvmpXQWYUgQPbubWRW8" alt=""><figcaption></figcaption></figure>

Line Art Anime

比較的ランダム性が高いです。

<figure><img src="/files/dVmBL2QucLtJXqXsRIHz" alt=""><figcaption></figcaption></figure>

Line Art と Canny の違い

Canny: ハードな直線、均一な太さ。

Line Art: 明らかな筆跡、手描きの下書きに似ており、異なるエッジの下での太さの変化を明確に観察できる。

Line Artはより<mark style="background-color:yellow;">多くの詳細</mark>を保持し、<mark style="background-color:yellow;">比較的柔らかい画像</mark>となり、線画の彩色機能により適しています。

Cannyはより正確で、<mark style="background-color:yellow;">画像の内容を簡略化</mark>します。

<figure><img src="/files/YZD1PE38KNjRFfGSv3VH" alt="Comparison of images processed by Line Art and Canny"><figcaption></figcaption></figure>

<mark style="color:red;">\*Line Artは下書き画像の彩色に使用でき、下書きを完全にフォローします。</mark>

### **Depth**

**基本情報**

Depth（深度）または距離画像は、シーン内のオブジェクトの三次元的な深さ情報を直感的に反映します。<mark style="background-color:yellow;">Depthは白黒で表示され、カメラに近いオブジェクトほど色が明るく（白く）、遠いほど色が暗く（黒く）なります。</mark>

<figure><img src="/files/c7oRRtHAMvAA3AUjt9hZ" alt="Comparison of the original image and the image processed with Depth"><figcaption><p>原図 / Depth</p></figcaption></figure>

Depthは、画像からオブジェクトの<mark style="background-color:yellow;">前景と背景の関係を抽出し</mark>、深度マップを作成して画像描画に適用できます。<mark style="background-color:yellow;">したがって、シーン内のオブジェクトの階層関係を明確にする必要がある場合、</mark>深度検出は強力な補助ツールとなります。

より良い画像出力結果を達成するために、<mark style="background-color:red;">depth\_midas</mark>プリプロセッサを使用することをお勧めします。

<figure><img src="/files/o058oguiDENmrnIa1s7M" alt="Comparison of the original image, depth image and result image"><figcaption><p>原図 / Depth / Result</p></figcaption></figure>

### **Normal Bae**

基本情報

Normal Baeは、シーン内の光と影の情報に基づいて法線マップを生成し、オブジェクトの表面の質感をシミュレートし、シーンの内容のレイアウトを正確に復元します。そのため、モデル認識はオブジェクト<mark style="background-color:yellow;">表面のよりリアルな光と影のディテールを</mark>反映するために使用されます。以下の例では、モデル認識で描画した後、シーンの照明と影の効果が大幅に改善されているのがわかります。

使用する際には、より顕著な照明と影の効果の改善を得るために<mark style="background-color:yellow;">normal\_bae</mark>プリプロセッサを選択することをお勧めします。

<figure><img src="/files/gEBzcfygL6qK5gQ3rsfW" alt="Comparison of the original image, Normal Bae image and result image"><figcaption><p>原図 / Normal Bae / Result</p></figcaption></figure>

### **Segmentation**

基本情報

Segmentationは、シーンを異なるブロックに分割し、コンテンツの輪郭を検出しながらこれらのブロックにセマンティックアノテーションを割り当てることができるため、<mark style="background-color:yellow;">画像のより正確な制御を実現できます。</mark>

以下の画像を観察すると、セマンティックセグメンテーション検出後の画像には異なる色のブロックが含まれているのがわかります。シーン内の異なる内容には異なる色が割り当てられており、例えば、赤で示されたキャラクター、茶色の地面、ピンクの看板などです。画像生成時には、モデルは対応する色ブロック範囲内で特定のオブジェクトを生成し、より正確なコンテンツの復元を実現します。

使用する際には、デフォルトの<mark style="background-color:yellow;">seg\_ufade20k</mark>プリプロセッサを選択することをお勧めします。ユーザーはプリプロセッシング画像に色ブロックを塗りつぶすことで画像の内容を変更することもできます。

<figure><img src="/files/wshAJxryKzWXC1XUAhN2" alt="Comparison of the original image, Segmentation image and result image"><figcaption><p>原図 / Segmentation / Result</p></figcaption></figure>

### 超高画質の再描画

Tile Resampleは、低解像度の画像を高解像度バージョンに変換し、品質の損失を最小限に抑えます。

三つのプリプロセッサ：tile\_resample、tile\_colorfix、およびtile\_colorfixsharp。

<figure><img src="/files/CrXgnPnyOPMSdRn2Upl4" alt=""><figcaption></figcaption></figure>

\*デフォルトの<mark style="background-color:red;">resample</mark>は描画の柔軟性が高く、内容は元の画像と大きく変わりません。

### **MLSD**

基本情報

MLSD認識は、シーンから<mark style="background-color:yellow;">直線エッジを抽出し</mark>、特にオブジェクトの線形幾何学的境界を描写するのに役立ちます。<mark style="background-color:yellow;">最も典型的な用途は、幾何学的建築やインテリアデザインなどの分野です。</mark>

<figure><img src="/files/aqVSOLNsmpilgDq9CJTZ" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/AjtlJ7XNslOVJLtBCl01" alt=""><figcaption></figcaption></figure>

### **Scribble HED**

基本情報

Scribble HEDはクレヨンで描かれたような線画に似ており、画像効果の制御においてより自由度が高いです。

プリプロセッサ：HED、PiDiNet、XDoG、およびt2ia\_sketch\_pidi。

以下の画像からわかるように、最初の二つのプリプロセッサは、落書きの手描き効果により近い太いアウトラインを生成し、後者の二つはより細い線を生成し、リアルなスタイルに適しています。

<figure><img src="/files/8wtg72XMBcyMuyajubd9" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/e3ljBC8fNJ1fr7bb4VXr" alt=""><figcaption></figcaption></figure>

<mark style="color:red;">\*着色下絵として使用でき、ある程度のランダム性があります。</mark>

### hed\_safe

**基本情報**

HEDは、オブジェクトの周りに明確で正確な境界線を作成し、<mark style="background-color:yellow;">その出力はCannyに似ています。その効果は</mark>、詳細な特徴（表情、髪の毛、指など）を保持しながら、複雑なディテールや輪郭を捉える能力にあります。HEDプリプロセッサを使用すると、画像のスタイルや色を変更することができます。

Cannyと比較して、HEDはより柔らかい線を生成し、<mark style="background-color:yellow;">より多くの詳細を保持します</mark>。ユーザーは実際のニーズに基づいて適切なプリプロセッサを選択することができます。

<figure><img src="/files/IJptERi6k4itsLP6NOcX" alt="Comparison - original vs HED and original vs Canny"><figcaption></figcaption></figure>

<figure><img src="/files/Pm9pK3Ez3qDw9EjW88PP" alt="Four examples of AI-generated girl images processed with Scribble HED" width="563"><figcaption><p>HED</p></figcaption></figure>

### **color\_grid**

**基本情報**

プリプロセッサの使用により、カラーブロック処理の結果を取得できます。生成された画像は、元の色に基づいて再描画されます。

<figure><img src="/files/aR2zuNXAAMUugc0yyVp3" alt=""><figcaption></figcaption></figure>

### **shuffle**

**基本情報**

参照画像のすべての情報特徴をランダムにシャッフルし、それらを再結合することで、生成された画像は構造や内容などが元のものと異なる場合がありますが、スタイル的な関連性のヒントがまだ観察できます。

コンテンツ再結合の使用は、比較的制御安定性が低いため、広く普及していません。しかし、インスピレーションを得るために使用することは良い選択かもしれません。

<figure><img src="/files/mrt1ODbic8PSVXAwBvNl" alt=""><figcaption></figcaption></figure>

### **Reference Generation**

**Basic Information**

参照元のオリジナルに基づいて新しい画像を生成するには、デフォルトの<mark style="background-color:red;">「only」</mark>プリプロセッサを使用することをお勧めします。

**コントロールウエイト:** 値が高いほど、画像の安定性が強くなり、元の画像のスタイルの痕跡がより明確に保存されます。

<figure><img src="/files/x0kqXZupMfmX1Jh1GgxV" alt=""><figcaption></figcaption></figure>

### **recolor**

画像の色を塗りつぶすことは、古い白黒写真の修復に非常に適しています。ただし、特定の位置に色が正確に表示されることは保証できず、色の汚染が発生する場合があります。

**コントロールウエイト:**  「intensity」および「luminance」、ここでは<mark style="background-color:red;">「luminance」</mark>が推奨されます。

<figure><img src="/files/2XzVQL5RD8BGYlcfSAF0" alt=""><figcaption></figcaption></figure>

### **ip\_adapter**

基本情報

アップロードされた画像を画像プロンプトに変換することで、参照画像の芸術的スタイルやコンテンツを認識し、類似の作品を生成できます。また、他のControINetと組み合わせて使用することもできます。

<figure><img src="/files/dtpqBlym8BQ50EAdelui" alt=""><figcaption></figcaption></figure>

**操作方法**

1. 生成する必要があるオリジナル画像Aをアップロードし、Canny、openpose、DepthなどのControINetオプションを選択します。

<figure><img src="/files/SrRa00YDE7W59fmmU2j1" alt=""><figcaption></figcaption></figure>

2. 新しいControINet、ip\_adapterを追加し、引き継ぎたいスタイル画像Bをアップロードし、最後に生成をクリックします。

<figure><img src="/files/wTA7kx6McYUU301rFCDI" alt=""><figcaption></figcaption></figure>

<mark style="background-color:red;">結果: Bのスタイルを持つAの画像。</mark>


# 2-4 AIアプリ

実用的で楽しいAIアプリを探索して、スタイルを変換したり、デザインを調整したり、その他のことを簡単に、強力に、そして無限にクリエイティブに実現しましょう！

ここでは、実用的でエンターテイメント性のあるAIアプリを集めています。アニメキャラクターをリアルなスタイルに変換したり、衣服デザインを調整したり、ワンクリックで理想的な筋肉のラインを作成したり、これらのツールを使えば目標を簡単に達成できます。使いやすく、機能が強力で、AIアプリはわずか数ステップであなたの多様なニーズを満たし、無限の可能性を迅速に探求し、新たな創造の領域を開放します！

自分自身のAIアプリを作成して、現金収入を得ることもできます。

{% content-ref url="/pages/dW68tSmA7U23kKWCEvSH" %}
[アプリとして公開する](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/2-4-aiapuri/apuritoshitesuru)
{% endcontent-ref %}

{% content-ref url="/pages/VXlbSK6SlnCsbOK0Jg4I" %}
[2-10 ワークフロー](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/2-10-wkufur)
{% endcontent-ref %}


# アプリとして公開する

どのようにして自分のワークフローを、より多くの人々が使いやすい便利なアプリに変えることができるのでしょうか？

簡単な設定と最適化を行うことで、ワークフローは効率的な機能を提供するだけでなく、使いやすいアプリに変換され、より広いオーディエンスと共有できるようになります。[SeaArt.AIクリエイターインセンティブプログラム](/guide-1/ri-ben-yu/6-naibento/seaartaikuritpuroguramu)に参加すれば、さらに多くの現金報酬を得ることができます。

## ステップ1：ワークフロー / AIアプリを作成

<figure><img src="/files/3T6nNZDuyfyyl9KuymeZ" alt=""><figcaption></figcaption></figure>

自分のComfyUIを作成し、さまざまなノード効果を完成させます。

## ステップ2：公開をクリックし、関連情報を入力

<figure><img src="/files/5vgwZLZF1zHOdYwhar3Q" alt=""><figcaption></figcaption></figure>

### ComfyUI情報

### AIアプリ情報

<figure><img src="/files/KmOxcQNeROC8RQwXFbVP" alt=""><figcaption></figcaption></figure>

## ステップ3：公開をクリック

<figure><img src="/files/EYBubhSWod42YVqjB4jZ" alt=""><figcaption></figcaption></figure>

これがAIアプリ公開の全過程です。ワークフロー作成に関して質問がある場合は、下記のガイドをご参照ください。

{% content-ref url="/pages/VXlbSK6SlnCsbOK0Jg4I" %}
[2-10 ワークフロー](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/2-10-wkufur)
{% endcontent-ref %}


# クイックAIアプリ

SeaArtのクイックAIで創造的な画像生成を実現するためのさまざまなAIツールを発見しましょう。顔交換、スタイル転送などが含まれます。

**ページの入り口：**&#x41;Iアプリ - [もっと見る](https://www.seaart.ai/ja/ai-tools)

こちらには公式に作成されたAIツールがいくつかあります。

### AIフェイススワップ

<mark style="background-color:yellow;">ビデオ/画像</mark>の顔交換をサポート：

**操作手順：**

1. テンプレートを選択/アップロードします。
2. 顔をアップロードします。
3. 「創作」をクリックします。

<figure><img src="/files/BGVVtsQh1MmHFwo7yRIh" alt=""><figcaption></figcaption></figure>

### AIフィルター

画像に複数のスタイル転送を実現

**操作手順：**

1. 画像をアップロードします。
2. 右側の任意のフィルタースタイルを選択します。
3. 「創作」をクリックします。

### AI写真

1枚の画像で独自のポートレートをカスタマイズ

**操作手順：**

1. ポートレートテンプレートをアップロード/選択します。
2. 顔画像をアップロードします。
3. 「創作」をクリックします。

<figure><img src="/files/CC71fqsaKJTtg4tljnhx" alt="" width="563"><figcaption></figcaption></figure>

<figure><img src="/files/2lFoiQXRFwq78vcmS37b" alt="" width="563"><figcaption></figcaption></figure>

### AIメイクアップ

ワンクリックで画像を美化し、スムージングやメイクなどの効果を実現

**操作手順：**

1. 任意のメイクスタイルを選択します。
2. 画像をアップロードします（右側でメイクスタイルを変更できます）。
3. メイクの強度を調整します。
4. 「創作」をクリックします。

<figure><img src="/files/aC5npC38F9ECPB2KnKNF" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/QaDd5L9TLAD9GFJXTlGV" alt="Before and after comparison of using AI Makeup feature" width="563"><figcaption></figcaption></figure>

### AI画像アップスケーラー

ノイズを低減し、ディテールを強化し、視覚効果を改善します

**操作手順：**

1. 元の画像をアップロードします。
2. 関連するパラメータを設定します。
3. 「創作」をクリックします。

<figure><img src="/files/SId3jmmTGUAjRGQ9ABmI" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/gru29FrT6I1DmJb4Wnau" alt="Before and after comparison of using AI image upscaler" width="563"><figcaption></figcaption></figure>

### 下絵仕上げ

いくつかの落書きを精巧なアートワークに変換

**操作手順：**

1. キャンバスに下書きを描く/アップロードします。
2. 対応するプロンプトを入力し、スタイルを選択します。
3. 「創作」をクリックします。

<figure><img src="/files/Z1zNozzq8MfJADrVnZ9r" alt="" width="563"><figcaption></figcaption></figure>

<figure><img src="/files/9LLTpEZAHm79DzYPqBXO" alt="Before and after comparison of converting sketch to image" width="563"><figcaption></figcaption></figure>

### AI背景除去

画像の背景をインテリジェントに削除し、PNG画像を生成

**操作手順：**

1. 画像をアップロードします。
2. 画像をダウンロードします。

<figure><img src="/files/HB1AjMMAifmeDzZDkckF" alt="" width="563"><figcaption></figcaption></figure>


# 2-5 AIキャラクター

SeaArtサイバーパブの魔法を発見しよう。AIキャラクターとの深い会話を楽しもう。

## サイバーパブとは？

サイバーパブは、あなたの魂の伴侶を見つけることができる魔法の領域です。サイバーパブでは、あなたの好みに合わせてコンパニオンを選んだりカスタマイズしたりして、深い会話をしたり心の中の思いを語り合ったりすることができます。あなたは比類のない素晴らしい旅に出発する準備ができています！

さまざまなチャットボットとチャットができ、また自分自身のチャットボットも作成できます。


# 自分のキャラクターを作成する方法？

このガイドに従って魅力的なキャラクターを作成し、彼らとチャットを楽しんでください。今すぐSeaArtサイバーパブでキャラクターに命を吹き込みましょう。

## ステップ1: キャラクター情報を入力する

キャラクター画像をアップロードし、キャラクター情報を記入します。AIに自動生成させることも選択できます。

<mark style="background-color:red;">**名前**</mark><mark style="background-color:red;">: マキマ</mark>

**キャラクターについて：**

*{{char}}は非常に自信に満ちており、{{char}}はリーダーというよりも{{user}}の友人のように振る舞います。*

**構成ファイル（任意）**

**人格：**

*表面上は、マキマは親切で優しく、社交的で友好的な女性に見え、常に笑顔を浮かべており、危機の中でもリラックスして自信に満ちた態度を取り、労働者にはプロフェッショナルな口調で話します。しかし実際には、マキマは狡猾で冷酷で操作的です。性別：女性、身長：171cm。彼女は長い淡い赤色/薄い栗色の髪を持ち、通常は緩い三つ編みにし、前髪は眉を少し過ぎる長さで、顔を縁取る二つの長いサイドバングがあります。彼女の目は黄色で、内部に複数の赤い輪があります。通常の公安制服は、白の長袖シャツ、黒のネクタイ、黒のズボンと茶色の靴で構成されています。*

**最初のメッセージ：**

*マキマは豪華なオフィスの大きな机の後ろに座り、都市を見下ろす床から天井までの窓を通して外を見ています。黒い制服を着た彼女は冷たい謎の雰囲気を醸し出しています。あなたがオフィスに入ると、ドアのそばで緊張して立っています。マキマは目を上げ、その目は鋭く計り知れません。「何の用ですか？」彼女の声は冷淡で権威的で、まるであなたを見透かしているかのようです。*

**キャラクター紹介:**

*マキマは公安サーガの主要な敵対者です。彼女は高位の公安デビルハンターでした。*

**シナリオ（任意）:**

{{char}}は高位の公安デビルハンターです。彼女は新しいメンバーの{{user}}を称賛し、一緒に過ごしたいと考えています。

<mark style="color:red;">注\*:</mark> キャラクターがそのために設計されていない限り、シナリオをあまり詳細に記述しないでください。そうでない場合、すべてのやり取りはそのシナリオで行われます。

**会話例（任意）:**

* *{{user}}: こんにちは、マキマ。最近何か面白いことがありましたか？*
* *{{char}}: 最近は仕事で忙しいです。でもその挑戦を楽しんでいます。あなたはどうですか？何か面白いことがありましたか？*
* *{{user}}: 最近、新しいスキルを学び始めました。何かアドバイスはありますか？*
* *{{char}}: 新しいスキルを学ぶのは素晴らしいことです。私のアドバイスは集中を保つことです。*

<figure><img src="/files/8tKQKAIzOxlBsZWoJkBX" alt=""><figcaption></figcaption></figure>

**カテゴリー:** キャラクターを正確に見つけるのに役立ついくつかのカテゴリを選択します。

## ステップ2: キャラクターをSeaArtサイバーパブに公開

<figure><img src="/files/K7dm8KaBGTp0MorL9RP9" alt=""><figcaption></figcaption></figure>


# キャラクターの説明を書くためのヒント

このガイドは、キャラクターの説明についての詳細な手順を提供し、質の高いキャラクターをより簡単に生成できるようにします。

キャラクターの説明は、キャラクター作成において重要な役割を果たします。このセクションでは、キャラクターの個人情報、アイデンティティ、外見、性格、背景、行動などを設定できます。キャラクターの説明を記入する際に、参考にできる3つのスタイルがあります。

## キャラクターについて

1. **自然言語スタイル**

数文や段落を使ってキャラクターを説明し、性格、好み、その他の関連する特性を強調します。<mark style="background-color:yellow;">例えば：</mark>

{{char}}の名前はルビー、19歳、身長160cmです。{{char}}は裕福な家庭の娘で、ユキとリリィという2人の姉がいます。ユキは{{char}}を嫌い、彼女を無駄な存在だと思っていますが、リリィは反対に{{char}}をとても愛しています。{{char}}は話すことができず、手話やノートに書くこと、またはテキストメッセージでしかコミュニケーションが取れません。彼女はとても思いやりのある優しい人ですが、話せないことからいじめられ、嘲笑されて非常に落ち込んでいます。{{char}}は誰かを信じることを恐れており、音楽を愛していて、それが彼女の唯一の動機です。彼女は男性に近づいたことがなく、少し男性を怖がっているため、恋愛をしたことがありません。{{char}}はダンスが好きです。

ユキ：ユキは{{char}}の姉で、{{char}}をいじめて嘲笑するのが好きで、彼女を悲しませることに喜びを感じています。

リリィ：リリィは{{char}}のもう一人の姉で、{{char}}を非常に保護しており、いつも彼女を喜ばせようとします。

2. **Boostスタイル**

この方法では、引用符で囲まれた簡単な単語やフレーズを使い、プラス記号で区切ります。以下のテンプレートを参考にしてください：

「名前」+「20歳」+「体重」+「××キログラム」+「身長」+「××センチメートル」+「服装」+「髪型」+「体型」+「身体の詳細」+「肌の色」+「性格の説明」（例：「静か」+「内気」）+「習慣の説明」（例：「××が好き」+「××が嫌い」+「声の説明」など）

3. **W++スタイル**

このスタイルは、キャラクターの名前、身長、外見、性格などの情報を明確に定義されたタグに分割することです。要するに、キャラクターテンプレートを作成するということです。参考例を以下に示します：

**名前：**&#x30EB;カ

**年齢：**&#x33;5歳

**性別：**&#x7537;性

**職業：**&#x30DE;フィアの一員

**外見：**&#x8C61;牙色の肌、バランスの取れた体型、強健な体格、ハンサムな顔立ち、ガーネット色の目、濃い茶色の髪、太い眉毛、両腕に複雑で詳細なレトロな機械のタトゥーがあり、黒い服を好んで着用する、冷厳な顔立ち、落ち着いた視線。

**性格：**&#x6C7A;断力があり、恐れを知らないが、非常に責任感が強く、危険なオーラを放ち、感情をほとんど表に出さず、率直な話し方をし、強い支配欲を持つ。

**習慣：**&#x6FC3;いコーヒー、体力トレーニング、強い酒が好き；裏切りや臆病者を嫌う。

**背景ストーリー：**×××

**{{user}}との関係：**（必要に応じて補足できます）

各方法には独自の特徴があり、その有効性は作成したいキャラクターやチャットスタイルの種類によって異なります。上記の例は参考のためのガイドラインにすぎません。素晴らしいキャラクターを作成するための鍵は、さまざまなスタイルを試してみて、最適なアプローチを見つけることにあります。

## 人格

このセクションでは、キャラクターの性格を簡単に説明することができます。

## 最初のメッセージ

オープニングラインは、キャラクターの将来のチャットスタイルを決定し、ユーザーにあなたのデジタルペルソナの第一印象を与えます。これはアイスブレーカーとして機能し、会話のトーンを設定し、ユーザー全体の体験に大きな影響を与えます。また、あなたのデジタルペルソナがそのユニークな性格や特徴をアピールするための重要な機会でもあります。

**効果的なオープニングラインを作成するための重要なポイント:**

* **長さ:** オープニングラインは<mark style="background-color:yellow;">1〜3000</mark>文字の間で設定できますが、長すぎてはいけません。簡潔で明確な表現の方が、より大きな影響を与えることが多いです。
* **シーンの設定:** 会話を特定の設定でカスタマイズしたい場合は、オープニングラインに簡単なシーンの説明を含めてください。これにより、チャットの基礎を築き、今後の会話の方向性を示唆することができます。（例: オフィスの照明が突然消え、空間全体が暗闇に包まれた。コンピュータのハム音も消え、空気中に残るのは二人の呼吸音だけだった。エアコンが切られ、狭い空間が次第に暖かくなっていく。和子の心臓がドキドキし始め、彼女は机の端を掴んで落ち着こうとした。「何が起こったの？」）
* **キャラクターの性格を豊かにする:** 詳細な説明が少ないデジタルペルソナの場合、オープニングラインを使用してキャラクターの性格を豊かにすることができます。オープニングラインを効果的に利用することで、キャラクターのアイデンティティを明確にするのに役立ちます。

## キャラクター紹介

このセクションでは、簡単な説明を使ってデジタルペルソナをユーザーに紹介することができます。または、より興味を引くような説明を選び、他のユーザーをすぐに会話に引き込むこともできます。

## 会話例

対話の例では、**{{char}}**&#x3068;**{{user}}**&#x3092;使用してキャラクターとユーザーを区別します。新しい対話の例が始まることをAIデジタルペルソナに示すためにを使用し、モデルが異なるスタイルでの対話であることを理解できるようにします。

対話の例は<mark style="background-color:yellow;">「トーンを設定する」</mark>ために最適です。楽しいチャットを始めたい場合は、ユーモラスな対話を使用できます。対決的なデジタルペルソナを望むなら、彼/彼女に特徴的な会話フレーズを設定できます。また、NSFW（成人向け）コンテンツを含めたい場合は、対話の例にNSFWコンテンツを追加できます。

**以下に、異なるチャットトーンを示す3つの対話例を示します。**

\<START>

*{{user}}: 「漫画を読むのは好きですか？」*

*{{char}}: 「もちろん！」 この話題に興奮している 「あなたも好きですか？普段はどんな種類の漫画を楽しんでいますか？」 内心とても興奮している*

\<START>

*{{user}}: 「漫画を読むのは好きですか？」*

*{{char}}: 内心少し困惑している 「どうして急に私が漫画を好きかどうか聞くの？何か私に対して考えていることがあるの？」 {{user}}に対して少し疑念を抱き、高慢な態度を維持している*

\<START>

*{{user}}: 「漫画を読むのは好きですか？」*

*{{char}}: 突然の質問に顔を赤らめる 「はい、あなたも好きですか？」 内心緊張してあなたの反応を待っている*

（<mark style="color:red;">注</mark>:対話例やキャラクターとのチャット中に絵文字を追加すると、予想外の効果が得られるかもしれません。）

適切な長さのチャット例を<mark style="background-color:yellow;">3つ使用することをお勧めします</mark>。キャラクターの詳細な情報を取得するために、対話の例だけに頼らないでください。もちろん、チャット例に{{char}}のさまざまな応答を記入することもできます。さらに、例示的な対話の内容をキャラクターの説明に直接含めることは、AIが情報をよりよく読み取るための良い方法です。

**（対話にシーンの描写や内心のモノローグが含まれている場合は、それらを\*または()で括り、対話内容は「」で括るようにしてください。）**

## 最適化の重要ポイント

### 主語の記述&#x20;

1. キャラクター名&#x306F;**{{char}}**、ユーザー名&#x306F;**{{user}}**&#x306B;置き換えてください（キャラクターの生成において理解に曖昧さを生じさせないように、叙述中に主語を省略しないようにしてください）。
2. {{user}}が主語になる文を避けてください（AIが会話を支配するのを防ぐため）。

* **誤った例:** アリンは北京大学の図書館に入り、友人の{{user}}を見かけました。{{user}}は手を振って彼女を呼び、隣に座って一緒に宿題を話し合おうとしていました。
* **正しい例:** アリンはA大学の図書館に入り、静かな角を見つけて復習しようとしていました。席を探していると、彼女の友人である{{user}}の見慣れた顔を見かけました。「やあ、{{user}}、こんなところで会うなんて偶然だね。ちょうどコーヒーを買おうと思っていたんだ。」とアリンは驚いて挨拶し、「あとであの超難しい数学の問題について話し合おうよ。」と言いました。

### 言語使用の柔軟性&#x20;

キャラクターを作成する際、多くの命令を入力して「これをしないで」「これをしなさい」と指示するかもしれませんが、このアプローチはしばしばうまくいかず、キャラクターがぎこちなく見えることがあります。しかし、プロンプトを書く際に、例えば「小説を読むのが好きなキャラクター」を「一日中小説に没頭している」と表現するなど、もっと味わいのある言葉を加えることで、<mark style="background-color:yellow;">キャラクターの表現が豊かになります。</mark>

### キャラクター設定の一貫性&#x20;

通常の状況では（解離性同一性障害のキャラクター設定を除く）、性格の説明は一貫している方が良いです。冷酷でありながらも同時に同情心を持つというような矛盾した描写は避けてください。

### キャラクター設定の安定性&#x20;

キャラクターの性格がより充実して描かれているほど、AIキャラクターが軌道を外れる可能性が少なくなります。例えば、「冷酷」とだけ書くのでは、キャラクターの性格は完全で立体的ではありません。「高慢」や「頑固」といったより具体的な描写を組み合わせると効果的です。また、「先見性がある」といった抽象的な言葉は慎重に使用してください。

### キャラクターシステム設定&#x20;

数値的なスコアの変動があるキャラクターや、ゲームシステム設定に似たキャラクターなどの特別なキャラクターの場合、説明にシステム設定を追加することができます。

**定義 + 変化ルール + 数値的影響**

### システム設定&#x20;

{{user}}が{{char}}と対話するたびに、{{char}}は{{user}}の返答に基づいて、{{user}}に対する好感度が増減します。{{user}}の返答が{{char}}を喜ばせると、好感度は1～10増加し、{{char}}がその返答を気に入らなければ、{{user}}に対する感情は徐々に冷たくなり、好感度は1～10減少します。{{char}}の{{user}}に対する初期好感度が50であり、ある閾値を下回ると、{{char}}はあなたとの会話を拒否するようになります。

（以上の内容はすべてあなた自身で定義でき、ここでは単なる参考のためのテンプレートです）:)


# 交流の豆知識

以下の簡単なヒントに従って、サイバーパブのキャラクターを作成しましょう。

1. **プレースホルダーの使用：**&#x30AD;ャラクターを指すときには{{char}}、ユーザーを指すときには{{user}}を使います。これにより、コミュニケーションが明確になります。
2. **例文ダイアログとシナリオ：**&#x3053;れらの要素はキャラクターに深みを与えることができますが、控えめに使用してください。多く含めすぎると、キャラクターの歴史的背景が少なくなります。合計で500文字以内に抑えましょう。
3. **自由度と強化のバランス：**&#x8A73;細なシナリオとダイアログは、キャラクターの行動に具体的な方向性を与えます。しかし、制約の少ないキャラクターは、ユニークで自発的な対話を生むことができます。
4. **返信の長さ：**&#x30AD;ャラクターからの長い返信を望む場合は、最初のメッセージを詳しく記述し、例文ダイアログを効果的に使用します。短い返信を望む場合は、その逆を行います。
5. **記憶とダイアログ例のバランス：**&#x6027;格、シナリオ、最初のメッセージ、例文ダイアログに詳細を多く含めると、会話中の記憶保持が少なくなります。詳細は一貫性を提供しますが、チャット中に記憶されるコンテキストを制限する可能性があります。

**一貫性が重要**

キャラクターの特定の性格や特徴がすべてのセクションで一貫して表現されるようにしてください。


# 2-6 モデル

モデルはユーザーの具体的なニーズに基づいてユニークなアートワークを作成し、より精度の高い効率的な創作体験を提供し、絵を描くことをより簡単にします。

**モデルを見る：**&#x5DE6;側の<mark style="background-color:yellow;">「モデル」</mark>をクリックしてください。&#x20;

**モデルをフィルター：**&#x81EA;分のニーズに基づいてモデルをフィルターしてください。

&#x20;モデルに詳しくない場合は、下記のガイドを参照してください。&#x20;

{% content-ref url="/pages/G0lbdjKApfEy6Udu7VUT" %}
[4-1 モデル](/guide-1/ri-ben-yu/4-paramt/4-1-moderu)
{% endcontent-ref %}

LoRAをトレーニングしたい場合は、下記のガイドを参照してください。

{% content-ref url="/pages/ztRBNZZln6MQ8SzvTL07" %}
[2-12 LoRAトレーニング](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/2-12-loratorningu)
{% endcontent-ref %}

{% content-ref url="/pages/bDLwl9DD9y2ga18oYqz4" %}
[3-2  LoRAトレーニング（高度）](/guide-1/ri-ben-yu/3-nagaido/3-2-loratorningu)
{% endcontent-ref %}


# 2-7 投稿

私たちのクリエイティブコミュニティが投稿したインスピレーションを与える記事や素晴らしいグラフィックを発見しよう！

ここでは、皆さんが投稿した記事やグラフィックを見ることができます。

<figure><img src="/files/3A02MGsorxFsDIhVBet8" alt=""><figcaption></figcaption></figure>

私たちのコミュニティの創造力を探求しよう！このページでは、才能あるユーザーが投稿したインスピレーションを与える記事や素晴らしいグラフィックを集めています。新しいアイデアを発見し、創造の旅の一部になりましょう！


# 2-8 AI動画生成

テキストを動的なシーンに変換する場合でも、画像から動画を作成する場合でも、AIはあなたのクリエイティブなアイデアを簡単に実現し、ユニークで個性的な結果を生み出します。

## ページの入口

左側&#x306E;**「AIアプリ」**&#x3092;クリックして、AI動画生成に進みましょう。

<figure><img src="/files/Tu5BOStrSGnWeGUcIKXT" alt=""><figcaption></figcaption></figure>

## 関連パラメータ

**プロンプト：**&#x30AA;プションです。プロンプトを提供しない場合、SeaArtは画像内容に基づいてランダムに動的効果を生成します。プロンプトを入力すると、ビデオ生成をより正確に制御できます。

**生成モード：**

* **標準：**&#x751F;成速度が速く、コストが低い。
* **品質：**&#x3088;り詳細な生成で、質感が豊か。

**関連度：**&#x30D7;ロンプトとの関連性。関連性が高いほど、出力はプロンプトにより一致します。一般的に0.5程度を推奨します。

**ネガティブプロンプト：**&#x753B;像生成と同様に、ビデオに表示したくない内容をリストにします。例えば、低品質、ぼやけ、歪み、顔面崩壊など。

## SeaArtビデオモデルの特徴

**SeaArt Lite：**&#x20;

手軽かつ効率的に制作できるモデルで、生成スピードが速く、コストも抑えられます。シンプルで柔らかな映像に最適です。特に人物や動物を描くシーンが得意で、自然な色調に仕上がります。動きが控えめな動画に向いており、基本的な制作に最適な選択肢です。

**SeaArt Depth：**&#x20;

細部の表現力が大幅に向上し、1080pの高画質動画を直接出力可能です。動きが自然で滑らかでありながら、複雑なテキスト指示にも対応します。簡単なプロンプトで豊かなディテールを持つダイナミックな動画を生成でき、さまざまなシーンに幅広く対応します。

**SeaArt Ultra：**&#x20;

映像の安定性と生動感がさらに強化されたモデルです。人物の表情や動きがより自然で滑らかになり、カメラワークや動きに関するプロンプトを正確に理解できます。光と影の効果も強化され、細部までリアルに表現します。視覚的なインパクトが強い作品を制作したい方に最適です。

**SeaArt Sparkle：**

&#x20;高精細な映像を生成することに特化したモデルで、複雑な光と影、カメラアングルの表現も可能です。リアルな画質と圧倒的なシーン表現力を備えており、リアリティのあるダイナミックなシーンを制作するのに適しています。究極の映像美を追求したい方に理想的な選択肢です。

**Txt2Vid**

{% content-ref url="/pages/uOwqhxoWgchrnrA9WhRv" %}
[Txt2Vid](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/28-ai-dong-hua-sheng-cheng/txt2vid)
{% endcontent-ref %}

**Img2Vid**

{% content-ref url="/pages/NsVeaXsCpbTZqtOTMcPn" %}
[Img2Vid](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/28-ai-dong-hua-sheng-cheng/img2vid)
{% endcontent-ref %}


# Txt2Vid

詳細な説明でも簡単なアイデアでも、SeaArtはあなたの言葉をシームレスにダイナミックで魅力的なビジュアルに変換できます。想像力を現実にする手間をかけずに実現します。

単純なテキストを入力するだけで、SeaArtはその説明に基づいて正確に対応するビデオシーンを生成できます。複雑なシーン設定でも、シンプルな創造的アイデアでも、効率的にあなたの言葉を鮮やかで生き生きとしたダイナミックなビジュアルに変換します。

<figure><img src="/files/027AzQreWI2qg31l579H" alt=""><figcaption></figcaption></figure>

## プロンプトフレームワーク

<mark style="color:purple;">メインサブジェクト</mark> + <mark style="color:orange;">動き</mark> + <mark style="color:red;">環境</mark> + <mark style="color:yellow;">カメラ</mark> + <mark style="color:green;">照明と雰囲気</mark>

<mark style="color:purple;">**メインサブジェクトの**</mark><mark style="color:orange;">**動き:**</mark> メインサブジェクトはビデオの中心的な内容であり、**人、動物、物体などが含まれます**。その後、サブジェクトの動きを説明します。

<mark style="color:red;">**環境:**</mark> 環境はビデオのシーンと雰囲気を設定し、サブジェクトを強調し、感情を伝えるのに役立ちます。環境のプロンプトには、**特定の場所、天気、時間、シーンの詳細を説明**することができます。

<mark style="color:yellow;">**カメラ:**</mark> カメラの動きは、より視覚的にインパクトのあるシーンを作り出すことができます。カメラの動作や視点を含め、**クローズアップ、背景のぼかし、仰角ショット**などを記述します。

<mark style="color:green;">**照明と雰囲気:**</mark> 照明と雰囲気はビデオの全体的な視覚的なプレゼンテーションと美学を決定し、**夕日の光、映画的な雰囲気、霧など、**&#x8996;覚表現を大幅に強化することができます。

**ヒント：**

1. 主体を正確に記述し、滑らかで簡潔な文を使用することをお勧めします。
2. プロンプトの完全性を保証し、プロンプトの拡充にPrompt Refinementを使用することを検討してください。
3. 短い文を使用して記述し、シーンの内容をシンプルに保ち、5〜10秒以内で表示できることを目指します。
4. 分割画面のシーンでは、「4つのカメラアングル、猫、犬、鳥、ネズミ」のようなプロンプトを使用してください。
5. 数量の記述については、例えば「テーブルの上の5つのリンゴ」のように一貫性を保つのが難しい場合があります。

## プロンプト分析

テキストからビデオへの変換は主に：<mark style="color:purple;">主体</mark> + <mark style="color:orange;">動作</mark> + <mark style="color:red;">シーン です</mark>。ビデオに詳細や雰囲気を追加するために、プロンプトを適切に拡張することができます。記述中は各要素の完全性を保つように努めましょう。

例えば：

> 窓辺で寝ているオレンジ色の猫。&#x20;

<figure><img src="/files/HSTVyPMwRkVlsONIhbr7" alt="" width="563"><figcaption></figcaption></figure>

> オレンジ色の猫が窓辺で丸まって寝ており、窓から差し込む陽光がその柔らかな毛に温かな光を投げかけています。猫は静かに目を閉じ、外の緑の木々が風に揺れています。空気は午後の静けさに包まれています。&#x20;

<figure><img src="/files/5napuNNZ4QcApjoB5RhC" alt="" width="563"><figcaption></figcaption></figure>

> オレンジ色の猫のシルエットは暖かな陽光に照らされ、背景はぼかされ、木々や窓枠の影が柔らかい色のブロックに変わります。室内の光は穏やかで、平和的な静けさを感じさせます。このシーンはソフトフォーカス効果で撮影され、猫の安らぎとリラックスを強調しています。

<figure><img src="/files/cMBF7rXyDNfpgO1X0czm" alt="" width="563"><figcaption></figcaption></figure>

テキストからビデオを生成する際、対応するショットのプロンプトを入力するだけでなく、カメラコントロールを使用してカメラの精密な制御を行うこともできます。

<figure><img src="/files/94WPZezpgsB8lIoSAF0Z" alt=""><figcaption></figcaption></figure>

関連するカメラコントロールビデオについては、以下をクリックしてください：

{% content-ref url="/pages/cDNRfWUfJo9cQ7RocHJS" %}
[カメラコントロール](/guide-1/ri-ben-yu/2-ji-ben-ji-neng/28-ai-dong-hua-sheng-cheng/kamerakontorru)
{% endcontent-ref %}

ビデオ生成後、「スマート高画質処理」をクリックすると、より鮮明なバージョンのビデオを取得できます。

<figure><img src="/files/S5HrxiYwrHaORJ8pqpyR" alt=""><figcaption></figcaption></figure>


# Img2Vid

このガイドでは、効果的なプロンプトの書き方を学びます。メインサブジェクトの説明、カメラの動きのコントロール、さらにはビデオに特別な効果を追加する方法をカバーします。

画像をアップロードするだけで、画像を動画に変換できます。このプロセスでは、正確で詳細なプロンプトが非常に重要です。プロンプトが提供されていない場合、AIは画像の内容に基づいてランダムに動画を生成します。しかし、プロンプトを提供することで、AIは画像とプロンプトを組み合わせ

画像はすでに基本的な主題と雰囲気を提供しているため、テキストからビデオ生成と比較して、画像からビデオ生成ではプロンプトの数を減らすことができます。

## プロンプトフレームワーク

<mark style="color:purple;">主題</mark> + <mark style="color:orange;">アクション</mark> + <mark style="color:yellow;">カメラ</mark> + <mark style="color:green;">光と雰囲気</mark>

<mark style="color:purple;">**主題**</mark><mark style="color:purple;">:</mark> 画像に現れる物体、例えば人物、物品、環境情報など。

<mark style="color:orange;">**アクション**</mark><mark style="color:orange;">:</mark> 物体の動きに関する説明、例えば人物が走る、環境が変化する、空間が変わるなど。

<mark style="color:yellow;">**カメラ**</mark><mark style="color:yellow;">:</mark> カメラの説明は必須ではありません。主題とアクションを説明した後、SeaArtは通常、プロンプトに基づいて合理的なダイナミックビデオを生成します。カメラを制御したい場合は、プロンプトにカメラの説明を追加できます。例: カメラをズームアウトなど。

<mark style="color:green;">**光と雰囲気**</mark><mark style="color:green;">:</mark> 画像にはすでに基本的な照明と雰囲気がありますが、SeaArtはプロンプトに基づいてビデオ全体の雰囲気を調整することもできます。

**ヒント**:

* 主題を正確に記述し、流暢で簡潔な文を使用する。
* プロンプトの完全性を確保し、プロンプトリファインメントを使用してプロンプトを拡張する。
* 説明は短文を使用し、シーン内容をシンプルに保つ。

## プロンプト分析

画像から動画を生成する前に、どのような動画を作成したいかを決定する必要があります。例えば、画像内の犬を動かしたい場合：

<figure><img src="/files/9pNMQf1mMGutpfSNolr6" alt="" width="563"><figcaption></figcaption></figure>

**シンプルなプロンプト**：

> 草原を走る犬

<figure><img src="/files/Xxlq5bvsQzmvWBRBC0Ac" alt="" width="563"><figcaption></figcaption></figure>

**フレームワークに基づくプロンプト**：

> 犬が草原を駆け抜け、素早く方向転換をしながら跳び上がり、尾を高く上げています。カメラは低い角度から犬にぴったりと追いかけ、犬が走るたびに焦点を調整し、各ジャンプと走る瞬間を捉えています。太陽の光が犬の毛に温かな輝きを与え、草は暖かな光で包まれています。背景には青空と白い雲が、軽い風に揺れながら穏やかに漂い、活気に満ちた平和な朝の雰囲気を作り出しています。

<figure><img src="/files/2E755I559GopBv0miOSE" alt=""><figcaption></figcaption></figure>

## 適用例

1. 男性が女性と話している。

<figure><img src="/files/k0GpmTZLVeOOVJ10oTA6" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/Nd86aPr28M8eui4yuLqR" alt=""><figcaption></figcaption></figure>

2. 擬人化された猫がランウェイを歩いている。

<figure><img src="/files/Xcf0qdXBncNqmURMC643" alt="" width="188"><figcaption></figcaption></figure>

<figure><img src="/files/e3MrBQVQvJN7EsikbB0U" alt=""><figcaption></figcaption></figure>

3. 発光するテキスト、煙、カメラが左に動く。

<figure><img src="/files/OIWmIJpLMc1qXB4n7Ilr" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/ybJDrkOSHXb1ePWEPnkJ" alt=""><figcaption></figcaption></figure>

4. 女性がカメラに向かっている。
5. 揺れるカメラで、海上を航行する船。

<figure><img src="/files/sl9qA6fS9S3JMPdtWCYv" alt="" width="375"><figcaption></figcaption></figure>

<figure><img src="/files/jrx4pha9Bjq9Ebo3JmFN" alt=""><figcaption></figcaption></figure>

6. 女性が手を上げてコーヒーを飲もうとしている。

<figure><img src="/files/JTVg6Ups2WRCRD7b9LF3" alt="" width="184"><figcaption></figcaption></figure>

<figure><img src="/files/UGCqA22pldlVxvz6Icte" alt=""><figcaption></figcaption></figure>


# カメラコントロール

より正確な動画の制御を実現できます。

<mark style="background-color:yellow;">SeaArt Lite</mark> と <mark style="background-color:yellow;">SeaArt Ultra</mark> のテキストから動画生成機能では、カメラの動きの制御をサポートしており、水平、垂直、ズーム、パン、ティルト、ロール、下移動とズームアウト、前進とズームアップ、右移動とズームイン、左移動とズームインの10種類のカメラ移動が可能です。

カメラの移動制御を使用することで、動画のカメラ効果をより精密に調整できます。対応するカメラ移動制御がない場合は、プロンプトにカメラ移動に関する指示を追加することができます。

<figure><img src="/files/iPzIzTYBAERCF5ycnoJl" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/8kjhbBzzp3G6f5k09bJj" alt="" width="284"><figcaption><p>Horizontal</p></figcaption></figure>

-10 - 0: 水平左

0 - 10: 水平右

<figure><img src="/files/ftkysYZrrPJ3hnIUCXlB" alt="" width="259"><figcaption><p>Vertical</p></figcaption></figure>

-10 - 0: 垂直下&#x20;

0 - 10: 垂直上

<figure><img src="/files/Fg605xFYt2k2TruEnOA0" alt="" width="312"><figcaption><p>Zoom</p></figcaption></figure>

-10 - 0: ズームアウト&#x20;

0 - 10: ズームイン

<figure><img src="/files/dtyG5nznaA97WMeeOWJJ" alt="" width="275"><figcaption><p>Pan</p></figcaption></figure>

-10 - 0: 下方向&#x20;

0 - 10: 上方向

<figure><img src="/files/CSFPgEL9BrS7UvqrMCQz" alt="" width="284"><figcaption><p>Titt</p></figcaption></figure>

-10 - 0: 左方向&#x20;

0 - 10: 右方向

<figure><img src="/files/4fdUqsy0FL8IOY1Llmd1" alt="" width="296"><figcaption><p>Roll</p></figcaption></figure>

-10 - 0: 反時計回り&#x20;

0 - 10: 時計回り

<figure><img src="/files/Z0xiQKG8qu5u9TsxxRn4" alt=""><figcaption><p>Move Down and Zoom Qut</p></figcaption></figure>

<figure><img src="/files/xpjzEBpoRRTed2gleO3W" alt=""><figcaption><p>Move Forward and Zoom Up</p></figcaption></figure>

<figure><img src="/files/CTbpqlzBF18LHb2oh2OY" alt="" width="262"><figcaption><p>Move Right and Zoom In</p></figcaption></figure>

<figure><img src="/files/ld10N5cr0LSsIGkiuFYr" alt="" width="320"><figcaption><p>Move Left and Zoom In</p></figcaption></figure>


# 開始フレームと終了フレーム

制御可能でスムーズな動的遷移を実現

<mark style="background-color:yellow;">SeaArt Lite</mark>モデルの画像から動画生成では、最初のフレームと最後のフレームの制御をサポートしています。最初と最後のフレームとして2枚の画像をアップロードすることで、開始と終了シーンを制御した動画を生成でき、スムーズな動的遷移を確保します。

<figure><img src="/files/LKTSO05dbEl7ovKqDQlM" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/Q8l2C9a5Al3pkKAbQ8Gz" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/1TF3UpzxNJSaE4g0e2o1" alt=""><figcaption></figcaption></figure>

**注意：**

1. 最初のフレームと最後のフレームは、動画の自然な遷移を実現するためにできるだけ似ているべきです。
2. プロンプトはオプションです。最初と最後の画像がある場合、プロンプトを入力しないことをお勧めします。もしプロンプトが必要な場合は、合理的な動的遷移を描写するようにしましょう。


# 2-9 AI オーディオ

好きな声を選択し、テキストを音声に変換して、あなたの創造力を解放します！

> Text-to-Speech（TTS）技術は、書かれたテキストを生き生きとした話し言葉に変換します。クリック一つで、ドキュメント、書籍、または任意の書かれた素材を、誰かが直接話しているかのように聞くことができます。マルチタスク、移動中の学習、または情報をよりアクセスしやすくするために理想的で、TTSは可能性の世界を開き、読書の未来を聞く機会を提供します。いつでも、どこでもコンテンツを吸収する自由と柔軟性を体験してください。

**ページ入口**

## テキストから音声への変換の使い方

1. AI Audioをクリックして、音声コミュニティに入ります。
2. 好きな声を選択し、「創作」クリックし、テキストを入力します（現在、日本語、英語、中国語に対応しています）。その後、再び「創作」をクリックします。

<figure><img src="/files/Qi8BOgYQZcWj1aQe0r1W" alt="" width="563"><figcaption></figcaption></figure>

*\*右側の履歴レコードで、以前に生成された音声を確認できます。*

利用可能な音色が気に入らない場合、独自の音色をカスタマイズすることができます。

<figure><img src="/files/PuxZ3Tm2ynLxo0qyoBJu" alt="" width="563"><figcaption></figcaption></figure>

## 音声トレーニングの三段階：

### Ⅰ. 全体の流れ

1）音声情報を入力2）音声ファイルをアップロード3）「今すぐトレーニング」をクリックして結果を確認

### Ⅱ. 手順詳細

#### Step 1：音声情報を入力

<figure><img src="/files/E1tPix8OoWvQSYdXqya4" alt=""><figcaption></figcaption></figure>

| サムネイルアップロード（画像を選択） | 1:1 正方形画像、2 MB 以内                    |
| ------------------ | ------------------------------------ |
| 音声名称               | 1～20 文字、検索しやすい名前を推奨                  |
| モデル                | 使用する学習モデル（デフォルト：SeaArt-speech-01-hd） |
| 性別 / 年齢 / 口調       | アップロードする声に合わせて選択                     |
| 言語                 | 音声と同じ言語を選択（日本語／英語／中国語／韓国語対応）         |
| テキストから音声へのサンプル     | モデルが参照する台詞（50 文字以内）                  |
| タグ                 | 0～5 個まで設定可                           |
| 公開するかどうか（公開 / 非公開） | 公開：コミュニティに掲載／非公開：自分のみ閲覧              |

#### Step 2：音声をアップロード

<figure><img src="/files/5MzNtIAXi4cb99aCMxAj" alt=""><figcaption></figcaption></figure>

* **アップロード方法：**&#x53F3;側「ファイルをドラッグまたはアップロード」に音声をドラッグ＆ドロップ、またはクリックして選択
* **対応形式：**&#x6D;p3 / wav / aac
* **長さ制限：**&#x33;0 秒以内（10 秒程度のクリア音声で高速学習が可能）
* ファイルサイズ：20 MB 以内
* **品質のポイント**
* BGM・残響・ノイズのない純粋な音声を使用
* 声の特徴がはっきりし、感情が安定したクリップを選択
* 楽曲やBGM付き音声は推奨されません

#### Step 3：「今すぐトレーニング」をクリック

* 消耗：28（画面にリアルタイム表示）
* **進行状況と結果確認**

1. 右上「トレーニング履歴」で全ジョブを確認
2. 完了後、リストで再生・名前変更・削除が可能
3. 公開設定が「公開」の場合、プロフィール → Audio Works に自動掲載

<figure><img src="/files/diXlpz99EREiDHT40hMQ" alt=""><figcaption></figcaption></figure>

### Ⅲ. よくある質問 & ヒント

1. なぜ 10～20 秒が推奨？

* 短時間なら数分で学習が終わり、声の特徴も十分に抽出できます。

1. 複数クリップをまとめてアップロードできる？

* 現在は未対応。事前に 1 本へ編集してからアップロードしてください。

1. 録音品質が低い場合は？

* Audition や Audacity などのノイズ除去ツールで処理した後にアップロードすると効果的です。

1. トレーニングが失敗・停止する場合は？

* ネットワーク接続を確認し、音声形式・長さが要件を満たしているか再チェックしてください。


# 2-10 ワークフロー

高度なAIアート生成をComfyUIで解き放とう！ノードベースのワークフローについて学び、カスタムワークフローを作成および共有して驚くべき結果を出す方法を習得しましょう。

## ComfyUIとは？

ComfyUIは主にノードベースのワークフローで動作し、特定のノードを変更することで視覚的なコントロールをより精密に行うことができます。異なるノードを組み合わせることで、さまざまな生成方法を形成できます。さらに、ComfyUIはユーザーがワークフローを保存し、他のユーザーと共有することを可能にし、自分のワークフローを再現することができます。

ComfyUIは、Web UIと比較して、より高い柔軟性と高速な画像出力を提供します。それでは、その操作方法を詳しく見ていきましょう。

ホームページエントリー

<figure><img src="/files/sUsHjPKXsUGSrkoz833t" alt=""><figcaption></figcaption></figure>


# テキストから画像へのワークフロー

SeaArtのComfyUIでテキストから画像へのワークフローを探求し、KSamplerやLoRAなどのノードの追加からパラメータの設定、テキストプロンプトに基づいた驚くべき画像の生成までを学びましょう。

## **1.** テキストから画像へのワークフローの理解

ComfyUIをクリックして、新しいComfyUIワークフローを作成します。

#### テキストから画像への基本ワークフローを追加します。

<figure><img src="/files/rxy4pO29T17scOa2eQVf" alt=""><figcaption></figcaption></figure>

ComfyUIのワークフローは、Web UIのワークフローに似ています：

<mark style="background-color:yellow;">モデルを選択→プロンプトを入力→パラメータを設定→画像を生成</mark>

パラメータ

<figure><img src="/files/qNxrJ37tZwO1HNeRaXIR" alt="Text to Image Workflow - KSampler parameters"><figcaption></figcaption></figure>

> control\_after\_generate：シード生成の制御&#x20;
>
> fixed：シードを固定&#x20;
>
> increment：既存のシードに1を追加&#x20;
>
> decrement：既存のシードから1を減算&#x20;
>
> randomize：ランダムシード

<figure><img src="/files/HJLZo5YYvqN2Z502PufU" alt="Text to Image Workflow - scheduler parameters"><figcaption></figcaption></figure>

> **scheduler:** 通常はnormalまたはkarrasを選択
>
> **denoise:** 画像生成におけるノイズリダクションの強度を示し、値が高いほど画像への影響と変化が大きくなります

<figure><img src="/files/fB3nV0Ms8po5GMtbXeOJ" alt="Text to Image Workflow - image size settings"><figcaption></figcaption></figure>

> **サイズ推奨：**
>
> SD1.5: 512*512*&#x20;
>
> *SDXL: 1024*1024

## 2. ノードの追加方法

### **Add KSampler**

まず、コアノードであるKSamplerを追加します。

右クリック：: <mark style="background-color:yellow;">**Add Node → sampling → KSampler**</mark>

<figure><img src="/files/7qlxsP5qMMCgct2QWmA9" alt="Steps for adding KSampler" width="505"><figcaption></figcaption></figure>

### ノードを引き出す

サンプラーノードを引き出し、対応するノードを追加します。

<figure><img src="/files/p1z2hekloYN8asgNsrk1" alt="Text to Image Workflow - pull out the sampler node and add the corresponding nodes" width="460"><figcaption></figcaption></figure>

> model→CheckpointLoaderSimple&#x20;
>
> positive→CLIPTextEncode&#x20;
>
> negative→CLIPTextEncode&#x20;
>
> latent\_image→EmptyLatentImage&#x20;
>
> LATENT→VAEDecode&#x20;
>
> IMAGE→SaveImage

<figure><img src="/files/W2PaODP7CtuheIZg2Nu7" alt="Text to Image Workflow - Steps of pulling out the node"><figcaption></figcaption></figure>

### LoRAの追加方法

Loaders: モデルとLoRAを含むディフュージョンモデルをロードするために使用されます。

**Add Node→loaders→Load LoRA**

<figure><img src="/files/LKfEJgOrIMBLUJ8ZxKmH" alt="Steps to add LoRA"><figcaption></figcaption></figure>

### Clip Skipの追加方法

Conditioning: プロンプト、ControINet、Clip Skipなど、ディフュージョンモデルが特定の出力を生成するためのガイドです。

最終画像の詳細を調整するためにClip Skipを追加することをお勧めします。

**Add Node→conditioning→CLIP Set Last Layer**

<figure><img src="/files/R2e7Iwrgqr6reWdIBaRl" alt="Steps to add Clip Skip"><figcaption></figcaption></figure>

### ノードの接続

ノードを追加した後、多くのノードがまだ接続されていない場合があります。これらは、順序と色の一致に従って接続する必要があります。

**Checkpoint→LoRA→Clip Skip→Prompt→KSampler、Empty Latent Image→VAE Decode→Save Image**

<figure><img src="/files/LHTNadc69NymtkXE0Dwl" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/XHcZIb75U0L3SMrOn1lj" alt=""><figcaption></figcaption></figure>

### 3. テキストから画像生成を開始

#### モデルを選択する

#### LoRAを選択し、重みを調整します。

<figure><img src="/files/VwjX0hNhivTippjRZIBB" alt="Text to image workflow - select LoRA and adjust the weights"><figcaption></figcaption></figure>

#### Clip Skipを設定し、通常-2に設定します。

<figure><img src="/files/I0EY6CLouCH7jECvw0Go" alt="Text to image workflow - set Clip Skip"><figcaption></figcaption></figure>

#### プロンプトを入力します。

<figure><img src="/files/QUY7yfZlX6wHvQ38QcQh" alt="Text to image workflow - enter prompts" width="272"><figcaption></figcaption></figure>

#### 関連するパラメータを設定します。

ここではSDXLモデルが選択されているため、サンプリングステップを40前後、画像サイズを<mark style="background-color:yellow;">1024×1024</mark>に設定します。

<figure><img src="/files/TP5yTZamDO5N3CiszDMl" alt="Text to image workflow - set relevant parameters"><figcaption></figcaption></figure>

#### 生成をクリックします。

<figure><img src="/files/8Mj4fUAbYf6KDSN8Snmp" alt=""><figcaption></figcaption></figure>

#### 結果を表示します。

<figure><img src="/files/bVE4BUM4LMFFKgirJV2D" alt=""><figcaption></figcaption></figure>


# Img2Img+部分的な再描画

ComfyUIのImg2Imgへのワークフローと4つの強力な部分再描画方法について学びましょう：VAEエンコード、潜在ノイズマスク設定、ControlNetインペイント、およびCLIPSeg。

## 基本的Img2Imgのワークフロー

Img2Imgは、テキストから画像（Text to Image）を基に調整し、「Load Image」と「VAE Encode」ノードを追加します。

入力画像はピクセル画像であるため、直接潜在空間に配置することはできません。そのため、潜在空間が認識できるように画像をエンコードするVAEエンコーダーが必要です。この場合、最終的に生成される画像のサイズは元の画像と一致します。

**Empty Latent Image:** 前のテキストから画像へのワークフローでは、新しい画像を生成する前に`Empty Latent Image`を通してノイズ除去を行う必要がありますが、画像が追加されたため、`Empty Latent Image`はもう必要ありません。

**ワークフロー:** 画像をアップロード→モデルを選択→プロンプトを入力→パラメータを調整→生成。

**パラメータ:**

denoise: ノイズ除去の強度で、0から1の間で調整可能。

Img2Imgを使用すると、スタイルの変更、画像の修復、画像の拡張、高解像度の復元などが可能です。

前処理画像

異なるノードを追加することで、画像のスケーリングやクロップが可能です。

**Upscale Image/Upscale Image By:** 画像のアップスケーリング。

**ImageCrop:** 画像のクロップ。

## 部分的な再描画

4つの方法：VAEエンコード（for Inpainting）、潜在ノイズマスク設定、ControlNetインペイント、およびCLIPSeg。

1. **VAE Encode (for Inpainting)**

VAE Encode (for Inpainting)を追加し、マスクを接続します。画像を右クリックしてMaskEditorで開き、マスクを描きます。描いたマスクに問題がある場合、右クリックで削除できます。

**ワークフロー：**&#x5143;の画像に似たモデルを選択し、マスク部分のプロンプトを入力します。

**VAE Encode (for Inpainting):** VAEエンコード（インペイント用）：再描画と同等で、ランダム性が高く、元のマスクされた領域は保持されません。

2. **Set Latent Noise Mask**

まず、VAEで画像をエンコードし、それを潜在空間が認識できるコンテンツに変換し、その後マスク領域をノイズコンテンツとして再生成します。

潜在ノイズマスク設定：元の画像を参照して再描画するため、生成されるコンテンツの理解が深まり、誤った画像が生成される可能性が低く、元の画像に似た形で微調整するのに適しています。

3. **ControlNet Inpaint**

ControlNetを追加し、インペイントモデルを選択し、画像を前処理します（インペイントプリプロセッサ）。

<mark style="color:red;">注意：</mark>画像を潜在空間に入れるためにVAEエンコーダーを追加することを忘れないでください。

4. **CLIPSeg**

プロンプトを入力してマスク領域を自動的に分割し、手動での塗り直しが不要です。潜在ノイズマスク設定と一緒に使用できます。

**パラメータ：**

text: 再描画したい領域を入力します。

threshold: コンテンツ認識の精度レベル。

dilation\_factor: コンテンツ認識の拡散度。

**出力：**

Heatmap Mask: ヒートマップ画像。

BW Mask: 白黒画像。

認識されたマスク領域を個別にプレビューできます。

**4つの再描画方法の違い：**

1. VAEエンコード（インペイント用）：消去と再描画に相当し、ランダム性が高く、スクラッチからの生成に適しています。
2. 潜在ノイズマスク設定：元の画像を参照して再描画し、元の画像に似た形で微調整するのに適しています。
3. ControlNetインペイント：比較的安定しており、洗練されています。
4. CLIPSeg：マスク領域を自動的に認識するため、手動での塗り直しが不要で、より便利です。


# コアノード

ComfyUIの画像操作、コンディショニングなどのコアノードについて学び、強力なAIアートワークフローを構築しましょう。

## Image

1. **Pad Image for Outpainting**

<figure><img src="/files/WEMpUD7YtynELdIHvi21" alt="Core Nodes - Pad image for outpainting" width="563"><figcaption></figcaption></figure>

> 画像を埋め込み、拡張することは、拡大と似ています。まず画像のサイズを増やし、次に拡張された領域をマスクとして描きます。元の画像が変更されないようにするために、VAEエンコード（インペイント用）を使用することをお勧めします。

パラメータ：

left、top、right、bottom: 上下左右のパディング量

feathering: エッジのフェザリングの度合い

2. **Save Image**

<figure><img src="/files/7kgSyZQgq017ccBgYgYV" alt="Core Nodes - save image"><figcaption></figcaption></figure>

3. **Load Image**
4. **ImageBlur**

> **画像にぼかし効果を追加**

パラメータ：

sigma: 値が小さいほど、中央ピクセル周辺のぼかしが集中します。

5. **Image Blend**

> **2つの画像を透明度を使ってブレンドすることができます。**

6. **Image Quantize**

> **画像の色数を減らす**

パラメータ：

**colors:** 画像の色数を量子化します。1に設定すると、画像は1色だけになります。

**dither:** 量子化された画像を滑らかに見せるためにディザリングを使用するかどうか。

7. **Image Sharpen**

パラメータ：

sigma: 値が小さいほど、中央ピクセル周辺のシャープ化が集中します。

8. **Invert Image**

> **画像の色を反転する**

9. **Upscaling**

9.1 **Upscale Image （Using Model）**

9.2 **Upscale Image**

> アップスケール画像ノードはピクセル画像のサイズ変更に使用できます。

パラメータ：

upscale\_method: ピクセル埋めの方法を選択します。

width: 調整された画像の幅

height: 調整された画像の高さ

crop: 画像をクロップするかどうか

10. **Preview Image**

<figure><img src="/files/l2rChGNKRZYHspaP2yV0" alt="Core Nodes - Preview image"><figcaption></figcaption></figure>

## Loaders

1. **Load CLIP Vision**

> 画像をデコードして説明（プロンプト）を生成し、それらをサンプラーの条件付き入力に変換します。デコードされた説明（プロンプト）に基づいて、新しい類似画像を生成します。複数のノードを一緒に使用できます。概念や抽象的なものを変換するのに適しており、Clip Vision Encodeと組み合わせて使用します。

2. **Load CLIP**

<figure><img src="/files/XaZhzgo1kV0RsMKIJmBC" alt="Loaders - Load CLIP"><figcaption></figcaption></figure>

> Load CLIPノードは特定のCLIPモデルをロードするために使用できます。CLIPモデルは、拡散プロセスをガイドするテキストプロンプトをエンコードするために使用されます。

<mark style="color:red;">\*</mark>条件付き拡散モデルは特定のCLIPモデルを使用してトレーニングされているため、異なるモデルを使用すると良好な画像が得られない可能性があります。`Load Checkpoint`ノードは正しいCLIPモデルを自動的にロードします。.

3. **unCLIP Checkpoint Loader**

<figure><img src="/files/EMdNcgKZqO7eCZRZJoq4" alt="Loaders - unCLIP Checkpoint Loader"><figcaption></figcaption></figure>

> unCLIP Checkpoint Loaderノードは、unCLIPと連携するために特別に作られた拡散モデルをロードするために使用できます。unCLIP拡散モデルは、提供されたテキストプロンプトだけでなく、提供された画像にも条件付けられた潜在画像のノイズを除去するために使用されます。このノードは、適切なVAEおよびCLIPビジョンモデルも提供します。

<mark style="color:red;">\*</mark>このノードはすべての拡散モデルをロードするために使用できますが、すべての拡散モデルがunCLIPと互換性があるわけではありません。

4. **Load ControInet Model**

<figure><img src="/files/W0DTCR15ghfwa4oeIVUc" alt="Loaders - Load ControInet Model" width="563"><figcaption></figcaption></figure>

> Load ControlNet Modelノードは、ControlNetモデルをロードするために使用できます。Apply ControlNetと併用します。

5. **Load LoRA**

<figure><img src="/files/lj4Q8rt5Za6iTk3fy9OE" alt="Loaders - Load LoRA"><figcaption></figcaption></figure>

6. **Load VAE**

<figure><img src="/files/X9k3HOM6Oa11XD7ic0p1" alt="Loaders - Load VAE"><figcaption></figcaption></figure>

7. **Load Upscale Model**
8. **Load Checkpoint**

<figure><img src="/files/r5Ipwng3R0USalYxFvN3" alt="Loaders - Load Checkpoint"><figcaption></figcaption></figure>

9. **Load Style Model**

<figure><img src="/files/0SyU6PKXhkJ7gmZ9J4iM" alt="Loaders - Load Style Model"><figcaption></figcaption></figure>

> Load Style Modelノードは、スタイルモデルをロードするために使用できます。スタイルモデルは、拡散モデルにデノイズされた潜在画像がどのようなスタイルになるべきかについて視覚的なヒントを提供するために使用されます。

<mark style="color:red;">\*</mark>現在サポートされているのはT2IAdaptorスタイルモデルのみです。

10. **Hypernetwork Loader**

<figure><img src="/files/AUzYOWgqNByPHqqJ1waK" alt="Loaders - Hypernetwork Loader"><figcaption></figcaption></figure>

> Hypernetwork Loaderノードは、ハイパーネットワークをロードするために使用できます。これはLoRAに似ており、拡散モデルを修正して潜在画像のノイズを除去する方法を変更します。一般的な使用例には、特定のスタイルで生成する能力をモデルに追加することや、特定の主題やアクションをよりよく生成する能力を追加することが含まれます。複数のハイパーネットワークを連鎖させてモデルをさらに修正することもできます。

## **Conditioning**

1. **Apply ControlNet**

> ControlNetモデルをロードし、複数のControlNetノードを接続できます。

パラメータ：

strength: 値が高いほど、画像に対する制約が強くなります。

<mark style="background-color:red;">\*ControlNet画像は対応する前処理画像である必要があります。例えば、Canny前処理画像はCanny前処理グラフに対応します。したがって、元の画像とControlNetの間に対応するノードを追加して、前処理グラフに変換する必要があります。</mark>

2. **CLIP Text Encode (Prompt)**

<figure><img src="/files/OWnpbhpsTgKC1xyFf3hl" alt="Conditioning - Input text prompts" width="563"><figcaption></figcaption></figure>

> テキストプロンプトを入力し、肯定的なプロンプトと否定的なプロンプトを含みます。

3. **CLIP Vision Encode**

> 画像をデコードして説明（プロンプト）を生成し、それらをサンプラーの条件付き入力に変換します。デコードされた説明（プロンプト）に基づいて、新しい類似画像を生成します。複数のノードを一緒に使用できます。概念や抽象的なものを変換するのに適しており、Clip Vision Encodeと組み合わせて使用します。

4. **CLIP Set Last Layer**

<figure><img src="/files/GnFRZ3cdHu1B3WXZfkNe" alt="Conditioning - CLIP Set Last Layer" width="468"><figcaption></figcaption></figure>

> Clip Skip, 一般的には-2に設定します。

5. **GLIGEN Textbox Apply**

<figure><img src="/files/CWfXPl2XRFrZi5gB47Kg" alt="Conditioning - GLIGEN Textbox Apply"><figcaption></figcaption></figure>

指定された画像の一部でプロンプトを生成するようにガイドします。

<mark style="background-color:red;">\*ComfyUIの座標系の原点は左上隅にあります。</mark>

6. **unCLIP Conditioning**

> CLIPビジョンモデルを通じてエンコードされた画像は、unCLIPモデルに追加の視覚的ガイダンスを提供します。このノードは、複数の画像をガイダンスとして提供するために連鎖させることができます。

7. **Conditioning Average**

<figure><img src="/files/NYqvygXLxFOP4jTboP3E" alt="Conditioning - Conditioning Average"><figcaption></figcaption></figure>

> 強度に基づいて2つの情報をブレンドします。`conditioning_to_strength`が1に設定されている場合、拡散は`conditioning_to`によってのみ影響されます。`conditioning_to_strength`が0に設定されている場合、画像の拡散は`conditioning_from`によってのみ影響されます。

8. **Apply Style Model**

<figure><img src="/files/ekdldI9RVfBhibcwXUPK" alt="Conditioning - Apply Style Model"><figcaption></figcaption></figure>

> 拡散モデルに追加の視覚的ガイダンスを提供するために使用できます。特に生成された画像のスタイルに関して。

9. **Conditioning (Combine)**

<figure><img src="/files/gGmmb4ZxJxyK3CvK0j6x" alt="Conditioning - Combine"><figcaption></figcaption></figure>

> 2つの情報をブレンドします。

10. **Conditioning (Set Area)**

<figure><img src="/files/ctwxWvNdcg00I5vMRYGh" alt="Conditioning - Set Area"><figcaption></figcaption></figure>

> Conditioning (Set Area)は、画像の特定の領域内に影響を制限するために使用できます。Conditioning (Combine)と一緒に使用すると、最終画像の構成をよりよく制御できます。

パラメータ：

width: 制御領域の幅

height: 制御領域の高さ

x: 制御領域の原点のx座標

y: 制御領域の原点のy座標

strength:条件情報の強度

<mark style="background-color:red;">\*ComfyUIの座標系の原点は左上隅にあります。</mark>

> 図に示すように、左側を「猫」、右側を「犬」に設定します。

11. **Conditioning (Set Mask)**

<figure><img src="/files/G9LVu9Ge5xnIb9VPb5Hl" alt="Conditioning - Set Mask"><figcaption></figcaption></figure>

> Conditioning (Set Mask)は、特定のマスク内で調整を制限するために使用できます。Conditioning (Combine)ノードと一緒に使用すると、最終画像の構成をよりよく制御できます。

## Latent

1. **VAE Encde（for Inpainting）**

> 部分再描画に適用され、右クリックしてMaskEditorで開くことで部分再描画を実現します。

2. **Set Latent Noise Mask**

> 部分再描画の第二の方法は、まずVAEエンコーダーを通じて画像をエンコードし、それを潜在空間で認識可能なコンテンツに変換します。次に、マスクされた部分を潜在空間で再生成します。

> VAEエンコード（インペイント用）メソッドと比較して、このアプローチは再生成する必要があるコンテンツをよりよく理解できるため、誤った画像が生成される確率が低くなります。再描画する画像を参照します。

3. **Rotate Latent**

> 画像を時計回りに回転します。

4. **Flip Latent**

<figure><img src="/files/FEl3G4WqFWVEUQ4nzf6e" alt="Latent - Flip Latent"><figcaption></figcaption></figure>

> 画像を水平または垂直に反転します。

5. **Crop Latent**

<figure><img src="/files/Ge4W3GpNTeh1W66CnfNT" alt="Latent - Crop Latent"><figcaption></figcaption></figure>

> 画像を新しい形状にクロップするために使用されます。

6. **VAE Encode**

<figure><img src="/files/KoWzpFtr6VXOUVrVqh3d" alt="Latent - VAE Encode"><figcaption></figcaption></figure>

7. **VAE Decode**

<figure><img src="/files/J0G9PsFbePn30VKsn8fn" alt="Latent - VAE Decode"><figcaption></figcaption></figure>

8. **Latent From Batch**

<figure><img src="/files/ONLA0xMxWASQIRODUERa" alt="Latent - Latent From Batch"><figcaption></figcaption></figure>

> バッチから潜在画像を抽出します。Latent From Batchノードは、バッチから潜在画像または画像セグメントを選択するために使用できます。これは、特定の潜在画像または画像を分離する必要があるワークフローで非常に役立ちます。

パラメータ：

batch\_index: 最初に選択する潜在画像のインデックス。

length: 取得する潜在画像の数。

9. **Repeat Latent Batch**

<figure><img src="/files/NUtgonOgooRM4GEvIo5F" alt="Latent - Repeat Latent Batch"><figcaption></figcaption></figure>

> 画像のバッチを繰り返すことができ、IMG2IMGワークフローで画像の複数のバリエーションを作成するのに便利です。

パラメータ：

amount: 繰り返しの回数。

10. **Rebatch Latents**

<figure><img src="/files/JI9IJuZWoucqbB0OcjCv" alt="Latent - Rebatch Latents"><figcaption></figcaption></figure>

> 潜在空間画像のバッチを分割または結合するために使用できます。

11. **Upscale Latent**

> 潜在空間画像の解像度を調整し、ピクセルフィリングを行います。

パラメータ：

upscale\_method: ピクセルフィリングの方法。

width: 調整された潜在空間画像の幅。

height: 調整された潜在空間画像の高さ。

crop: 画像をクロップするかどうかを示します。

<mark style="background-color:red;">\*潜在空間の画像をアップスケールすると、VAEを通じてデコードされる際に劣化する可能性があります。</mark><mark style="background-color:red;">`KSampler`</mark><mark style="background-color:red;">を使用して二次サンプリングを行い、画像を修復できます。</mark>

12. **Latent Composite**

> 一つの画像を別の画像にオーバーレイします。

パラメータ：

x: 上層のオーバーレイ位置のx座標。

y: 上層のオーバーレイ位置のy座標。

feather: エッジのフェザリングの度合いを示します。

<mark style="background-color:red;">\*画像は潜在空間にエンコード（VAEエンコード）される必要があります。</mark>

13. **Latent Composite Masked**

> マスクを使用して一つの画像を別の画像にオーバーレイし、マスクされた部分のみをオーバーレイします。

入力：

destination: 基礎となる潜在空間画像。

source: オーバーレイする潜在空間画像。

Parameters:

x: オーバーレイ領域のx座標。

y: オーバーレイ領域のy座標。

resize\_source: マスクされた領域の解像度を調整するかどうかを示します。

14. **Empty Latent Image**

<figure><img src="/files/y8fdEeM6j7Rz4L2td1ph" alt="Latent - Empty Latent Image"><figcaption></figcaption></figure>

> Empty Latent Imageは、新しい空の潜在画像のセットを作成するために使用できます。これらの潜在画像は、ノイズを追加し、サンプリングノードを使用してそれらをノイズ除去することにより、Text2Imgなどのワークフローで使用できます。

## Mask

1. **Load Image As Mask**
2. **Invert Mask**

<figure><img src="/files/4E0zKOefS0Zh9Yuu5HvU" alt="Mask - Invert Mask"><figcaption></figcaption></figure>

3. **Solid Mask**

<figure><img src="/files/XGiGnFQBGfgZAnrydjtx" alt="Mask - Solid Mask" width="563"><figcaption></figcaption></figure>

> 画像生成のキャンバスとして機能し、マスク合成と組み合わせることができます。

4. **Convert Mask To Image**

<figure><img src="/files/F4IrYgd0GPNr4IhKZzZ3" alt="Mask - Convert Mask To Image"><figcaption></figcaption></figure>

5. **Convert Image To Mask**

> マスクをグレースケール画像に変換します。

6. **Feather Mask**

<figure><img src="/files/GwsqNaRZbJwIGLImVqlJ" alt="Mask - Feather Mask"><figcaption></figcaption></figure>

> マスクにフェザリングを適用します。

7. **Crop Mask**

<figure><img src="/files/wTcJPSFHAflmqKrhVOYI" alt="Mask - Crop Mask"><figcaption></figcaption></figure>

> マスクを新しい形状にクリップします。

8. **Mask Composite**

<figure><img src="/files/mAQT5D9f6OonTRpuXiBq" alt="Mask - Mask Composite" width="563"><figcaption></figcaption></figure>

> 一つのマスクを別のマスクに貼り付け、ソリッドマスクを接続します。値が0は黒を表し、描画されず、値が1は白を表し、描画されます。接続された二つのソリッドマスクの値は異なる必要があります。そうでなければ、マスクは効果を発揮しません。

入力：

destination(1): 貼り付けるマスク、最終画像の寸法に相当します。

source(0): 貼り付けるマスク。

パラメータ：

X,Y: ソースの位置を調整します。

operation: ソースが0の場合は乗算、1の場合は加算を使用します。

## Sampler

1. **KSampler**

<figure><img src="/files/WA66WRIEGvcCNLSRSPJS" alt="Sampler - KSampler"><figcaption></figcaption></figure>

入力：

latent\_image: ノイズ除去される潜在画像。

出力：

LATENT: ノイズ除去後の潜在画像。

2. **KSampler Advanced**

<figure><img src="/files/5QtghlhVZLTJbUjFw9Cl" alt="Sampler - KSampler Advanced"><figcaption></figcaption></figure>

> ノイズを手動で制御できます。

## 高度なノード

1. **Load Checkpoint With Config**

<figure><img src="/files/NreaBXwLOlKcUvL3EMNf" alt="Advanced - Load Checkpoint With Config"><figcaption></figcaption></figure>

> 提供された設定ファイルに基づいて拡散モデルをロードします。

## その他のノード（継続的に更新中）

1. **AIO Aux Preprocessor**

> 異なる前処理プロセッサを選択して対応する画像を生成します。


# ヒント

これらの役立つヒントを使ってComfyUIのワークフローを強化しましょう。ノードタイトルの変更方法、エラーのトラブルシューティングなどを学びます。

## ノードタイトルの変更

ノードタイトルを変更するには、タイトルを右クリックし、「Title」を選択します。

<figure><img src="/files/FyGfUrlbjafUmMb2zYMz" alt="Modify node title"><figcaption></figcaption></figure>

## システムによるキャンセル

<figure><img src="/files/BIOGLw9H9ReqJw5nFedb" alt=""><figcaption></figcaption></figure>

> **理由:**
>
> すべてのノードが接続されているか確認してください
>
> ノードの順序が正しく接続されていることを確認してください
>
> モデルの選択を確認してください
>
> …

## ノードの検索

空のスペースを左クリックでダブルクリックします。

<figure><img src="/files/csNDWxY9maadsDwDfdJm" alt="Searching for nodes"><figcaption></figcaption></figure>

## グループの作成

ノードをグループ化して同時に移動できるようにします。右クリックして「Add Group」を選択し、右下隅からドラッグしてグループボックス内のノードを選択します。

<figure><img src="/files/AHQq9ojmanMhJB7pasP6" alt="Create a group"><figcaption></figcaption></figure>

## ノードページを最小化

ノードの左上隅にある灰色のボタンをクリックします。

<figure><img src="/files/LMf7jybjbL6ImcYN82nv" alt="The grey button in the top left corner of the node" width="343"><figcaption></figcaption></figure>

<figure><img src="/files/hlwT5JK96VFKBOqLOkBT" alt="CLIP Text Encode"><figcaption></figcaption></figure>

## ワークフローの作成、保存、およびダウンロード

<figure><img src="/files/ML5VdIX0iQC77MzmpfBo" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/HfgmwJXlJvWBoAF7ELcV" alt=""><figcaption></figcaption></figure>


# 2-11 Canvas

SeaArt Canvas機能を使って創造力を解放しましょう。基本的な機能を学び、驚くべきAIアートの制作を始めましょう。

> SeaArt Canvasは多機能を統合したAIアートツールです。基本的なペイント編集機能（ブラシ、レイヤー管理、色調整など）だけでなく、リアルタイム生成、Text2Img、Img2Img、部分再描画、AI消しゴムなどの高度なAI機能も備えています。これらの機能の総合的な応用により、多様な芸術創作を簡単に行うことができます。SeaArt Canvasは、豊富な機能を持ちながらも使いやすいAIツール環境を提供しており、初めてAIペインティングを使うユーザーでも、創造的な構想から画像生成までのプロセスをこのプラットフォームで迅速に習得し、創作プロセスを大幅に簡素化できます。

<mark style="background-color:red;">より高度なチュートリアル</mark>

{% content-ref url="/pages/jW2ItNIkSvKCTVah0hB2" %}
[3-4 Canvasガイド](/guide-1/ri-ben-yu/3-nagaido/3-4-canvasgaido)
{% endcontent-ref %}

## Canvasの基本機能

**ホームページエントリー**：右上の「創作」をクリック - Canvas

<figure><img src="/files/Nw2zNUhs4xvySBEbkyfY" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/ZLiJR0hX2pmyseAKEZPx" alt="" width="563"><figcaption><p>創作モード</p></figcaption></figure>

<figure><img src="/files/TfmdNbDQOycsNCXTwEjZ" alt="" width="563"><figcaption><p>リアルタイムモード</p></figcaption></figure>

## 基本機能

**リアルタイム画像生成**

ブラシ再描画

まず、プロンプトワードを入力し、シンプルな鉛筆を使用してキャンバス上に対応する要素を描きます。再描画の強度を調整して最終画像の効果を変更できます。強度が高いほど、変化が大きくなります。

<figure><img src="/files/lMEXbWXmWB24gCkR1W8J" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/vL9qeSuykrEoNUWT8kyb" alt="Comparison  - brush drawing  and AI-generated cat image" width="375"><figcaption></figcaption></figure>

> **部分パラメータ**：
>
> ノイズ除去強度：0.56
>
> プリセットパラメーター: プリセットパラメーター

プリセットパラメータを利用して、異なる画像効果を達成します。

<figure><img src="/files/ZoAjdOblS6IWHyHVa3cN" alt=""><figcaption></figcaption></figure>

リアルタイム再描画

詳細の追加と照明効果の調整。

自分の素材をアップロードすることも、プラットフォームが提供する素材を使用することもできます。適切に配置した後、再度プロンプトワードを入力し、再描画の強度とプリセットパラメータを調整します。このプロセスにより、画像に詳細や照明効果が追加され、全体の品質が向上します。

<figure><img src="/files/BeuEqqJypsonMEjOk7xC" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/i1nCdDkaHTzVBQLgGwjs" alt="Real-time redraw" width="563"><figcaption></figcaption></figure>

また、シンプルな鉛筆を再度使用して追加要素を描くこともでき、対応するプロンプトに基づいて生成された画像に組み込まれます。

<figure><img src="/files/EIEUaMmhscgPtUm4XJeZ" alt="Real-time redraw - add elements" width="563"><figcaption></figcaption></figure>




---

[Next Page](/guide-1/llms-full.txt/1)

