# Introduction

Introduce Float16.cloud

Welcome to Float16.cloud documentation. We are a GPU managed service provider offering API-first solutions and resources for developers, with a focus on those new to the Large Language Model (LLM) industry who may lack sufficient resources for experimentation or implementation.

Our GPU managed service is a cloud-based solution that provides on-demand access to GPU resources. We manage the complexities of GPU infrastructure, including deployment, scaling, and maintenance. This allows our users to concentrate on their primary tasks, such as AI development and LLM applications, without concerns about hardware management.

As of April 2025, we are excited to introduce our new Serverless GPU Service. This service enables users to run, train, or deploy AI models and Python code on our high-performance GPUs, including H100 instances, without the need for complex configuration. Our pay-per-compute model ensures that users are billed only for the actual compute time used, offering a cost-effective and flexible solution. With this service, developers can focus entirely on their code without worrying about infrastructure setup and management.

See our quick start here.

{% content-ref url="/pages/TAehzihhRsI5OpDA4tnb" %}
[Serverless GPU](/getting-started/serverless-gpu)
{% endcontent-ref %}

This documentation will guide you through our services and help you make the most of our platform.

## Last Updated

{% hint style="info" %}
App Version 0.5.13 (23 June 2025)

* Added support for project names in create, list, and edit operations.
* Removing the server type task from deployment mode.

CLI Version 0.1.5 (June 2025)

* Support for project names in create, list, and edit operations.
  {% endhint %}

## How to Get Started?

To begin using our services, you need to sign up for a Float16 App account. For more details on this process, please refer to the [Account](/getting-started/account) section in our documentation. Once you have an account, you can start exploring and utilizing our range of services.

For organizations with sensitive data who prefer to host our service on their own infrastructure, we offer a private hosting option. Please contact our team for more information about this solution.

## Playground for developers

We offer a convenient Chat Playground for developers who want to test or quickly try out new models without implementing a user interface. This feature allows you to interact with our AI models directly and effortlessly.

Check out the playground,

{% content-ref url="/pages/rFGROrIRevotePWeXp6t" %}
[FloatChat](/getting-started/playground/floatchat)
{% endcontent-ref %}

We are excited to introduce a new playground designed for both developers and non-developers who wish to create and test few-shot prompts. This platform allows users to share their prompts with the community, offering free access to the SeaLLM-7b-v3 model and OpenAI API integration.

{% content-ref url="/pages/kpGqwElCUgRSzhQ7ZvXQ" %}
[FloatPrompt](/getting-started/playground/floatprompt)
{% endcontent-ref %}

We offers "Quantize by Float16" for developer who wants to compare inference speeds of leading LLMs: Llama, Gemma, RecurrentGemma, and Mamba. Test various quantization techniques and KV cache settings to help you optimize your LLM deployments and understand performance trade-offs.&#x20;

{% content-ref url="/pages/6wuuzLUQib3gjWxcgXaa" %}
[Quantize by Float16](/getting-started/playground/quantize-by-float16)
{% endcontent-ref %}

## Support

Our focus is on creating a developer-first community. We are excited to support you and the entire developer community. If you have any questions or issues with deploying, implementing, or launching your AI application, you can contact us or share your problem on [Discord](https://discord.gg/j2DVTMjr67). Additionally, if you have any requests or feedback you'd like us to know about, we're eager to hear from you.


# Account

Learn how to create your Float16.cloud account

<figure><img src="/files/D8a6Mtjfs4uQiw6TM6Hi" alt=""><figcaption><p>Sign up &#x26; Sign in</p></figcaption></figure>

## Sign up

When you first enter our web application, if you don't have an account yet, you'll need to sign up. We offer three ways to create an account:

1. Sign up with Google
2. Sign up with GitHub
3. Sign up with your email

For users choosing to sign up via a third-party provider (Google or GitHub), you'll need to accept the permission request for accessing your data from the provider. Upon successful sign-up, you'll be automatically logged into our console.

For users signing up via email, you'll need to set up your password. We'll then send a confirmation email to your provided email address.

## Activate Account

For users who sign up by email,

1. Check your email inbox for a confirmation email from Float16.cloud.
2. Click the confirmation link in the email to activate your account.
3. After clicking the link, your account will be activated, and you'll be automatically logged into our console.

## Sign in

If you already have an account, you can log in using Google, GitHub or your email and password.

Note:

* If you signed up via email using a Gmail address, you can also sign in with Google. Our system will recognize it as the same account.
* If you've forgotten your password, you can reset it by clicking the "Forgot Password" link.

***

Once you've entered the console, you can explore more features as detailed below.

## Let's Explore

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Dashboard</strong></td><td>View a dashboard with information about activities on your account</td><td><a href="/pages/7vf1oxNV1MPZzvPd3fYK">/pages/7vf1oxNV1MPZzvPd3fYK</a></td></tr><tr><td><strong>Payment</strong></td><td>Top-up and view current charges</td><td><a href="/pages/Xvl4yRXnhZyGxafJgYU0">/pages/Xvl4yRXnhZyGxafJgYU0</a></td></tr><tr><td><strong>Workspace</strong></td><td>Manage your workspace settings and configurations</td><td><a href="/pages/TPawPBC8naIL97lxcB9W">/pages/TPawPBC8naIL97lxcB9W</a></td></tr><tr><td><strong>Profile</strong></td><td>Review and update your personal profile information</td><td><a href="/pages/CqFMtEWEDda3Bd5Km1Jx">/pages/CqFMtEWEDda3Bd5Km1Jx</a></td></tr><tr><td><strong>Service Quota</strong></td><td>Estimate your current quota usage</td><td><a href="/pages/Y93Q2dyOqSFCEtdG92xH">/pages/Y93Q2dyOqSFCEtdG92xH</a></td></tr></tbody></table>


# Dashboard

View a dashboard with information about activities on your account

Our application features the Activity Dashboard. These dashboards provide users with comprehensive insights into their service usage and activities.

## Activity Dashboard

<figure><img src="/files/jITolLRHW6BsnksgmRH7" alt=""><figcaption><p>Activity Dashboard</p></figcaption></figure>

Provides detailed insights into user activities for each individual service, filtered by month (default: current month, up to 12 months of historical data available).

### Location

* Accessible within the Services, appears as a side menu "Activity".&#x20;

### Features

#### LLM as a Service

The dashboard includes four main components:

1. **Prompt Usage**: Daily prompt token usage across all models
2. **Completion Usage**: Daily completion token usage across all models
3. **API Requests**: Daily REST requests across all models

{% hint style="info" %}
Bars represent daily usage, with colors distinguishing between models.
{% endhint %}

4. **Activity Log**: Tabular view of all activities across all models, grouped hourly. Includes timestamps, requests, prompt usage, completion usage and average time for comprehensive usage tracking and analysis.

#### One Click Deploy

The activity dashboard appears in each instance, displaying individual instance activity logs. The dashboard comprises three main components:

1. **Token Usage**: Visualizes daily token consumption for this model instance. Each bar represents daily usage, with distinct colors differentiating between prompt and completion token usage.
2. **API Requests**: Illustrates daily REST requests made to this model.
3. **Activity Log**: Tabular view of all activities happen in this model instance, grouped hourly. Includes timestamps, requests, prompt usage, completion usage and average time for comprehensive usage tracking and analysis.

You can download your activity log as a CSV file. This log is organized hourly and separated by model. To access this feature, simply click the "Download CSV" button in the top of Activity Dashboard page.

#### Serverless GPU

The Activity dashboard appears in each project's activity section and provides a visual summary of usage and cost metrics. It consists of **four main components**:

1. **Total Cost:** Displays the total accumulated cost during the selected time range. The cost is categorized into three types: On-Demand, Spot and Storage.
2. **Total Usage:** Shows the total GPU compute time consumed, measured in seconds. Separates usage into On-Demand and Spot
3. **Cost Over Time (Graph):** A bar chart that visualizes daily cost over the selected month. Each bar is broken down by cost type: On-Demand, Spot, and Storage
4. **Tasks (Graph):** A daily breakdown of the number of tasks run within the selected period, categorized by Manual, Function, Server, and Spot triggers


# Profile

Review and update your personal profile information

## Profile Display

* Your profile picture is typically displayed in the top-right corner of the page or on the right side of the navigation bar.
* Clicking on your profile picture will show a dropdown menu containing:&#x20;
  * Your display name and email address
  * Settings
  * Log Out option

## Changing Your Display Name

<figure><img src="/files/S5hcJvSDwJFla6aQGi8R" alt=""><figcaption><p>Profile Setting</p></figcaption></figure>

1. Access "Setting" by clicking the app logo, your account<img src="/files/pz7AqJU7SjUmXJ2ribdo" alt="" data-size="line">in the bottom-left corner, or your profile in the top-right corner.
2. In the Settings menu, go to the "Profile" section. You'll see your registered email address and current display name.
3. Locate the "Name" field, enter your desired display name.
4. Click the "Save" button. A notification saying "Profile Updated" will appear, confirming the successful name change.

{% hint style="warning" %}

* Profile picture editing is currently not available.
* Your email address cannot be edited.
  {% endhint %}


# Payment

View current charges

## Overview

<figure><img src="/files/uvBlYR9r4VyY9hTeHYli" alt=""><figcaption><p>Payment Setting</p></figcaption></figure>

Our service uses a monthly billing cycle. By default, each account is assigned a spending limit of **$100**.

* On the **1st day of each month**, the system automatically generates a bill via Stripe Billing.
* You have **15 days** to make the payment before your account is **suspended**.
* After billing is generated, your spending limit will be reset to your original limit (e.g. $100 or your custom limit).

If you would like to increase your spending limit, please contact us.

> 📌 Once increased, the new limit will be permanent.

{% hint style="info" %}
For organizations interested in unlimited usage, feel free to [contact our team](https://email:support@float16.cloud/). We’re happy to provide customized payment plans tailored to your needs.
{% endhint %}

### Daily Credit

We are currently running a campaign that provides **$5 of free daily credit** to every account. The credit **expires at the end of each day**.

* To receive your daily credit, simply click the **“Claim Daily Free Credit”** button.
* Once successfully claimed, your **credit balance** will increase by $5.
* The campaign starts on **18 June 2025** and will continue until further notice.

{% hint style="warning" %}
Accounts that already have granted credits are **not eligible** for this campaign.
{% endhint %}

### Credit Usage

You can view your credit usage in the following locations

* Payment Settings
* Playground page of each project

Displayed information includes

* **Available Credit**: The total of your remaining granted and daily credits.
* **Daily Usage**: How much credit you’ve used today.

{% hint style="success" %}
Available credit will automatically update when credits expire or when a new billing cycle deducts your balance.
{% endhint %}

## Invoice

<figure><img src="/files/9iVhL2Z845UxCxUli9am" alt=""><figcaption><p>Invoice History</p></figcaption></figure>

To pay your invoice:

1. Click the app logo, account icon<img src="/files/pz7AqJU7SjUmXJ2ribdo" alt="" data-size="line">(bottom-left), or profile (top-right), then select "Settings" > "Payment".
2. Click the Invoice tab to view a list of monthly invoices.
3. Select the invoice you wish to pay. You'll be redirected to Stripe, our payment gateway. Follow on-screen instructions to complete payment.
4. Upon successful payment, you’ll be redirected back to the Payment page.

## Monitor your balance&#x20;

### Quick View

Estimated remaining spending credit is visible in the Payment settings overview tab.&#x20;

### Detailed Usage

<figure><img src="/files/zw5Pgn5W4qHYlYathN5t" alt=""><figcaption><p>Balance History</p></figcaption></figure>

1. Navigate to "Settings"
2. Select "Payment"
3. Click on the "Balance History" tab

On this page, you'll see a comprehensive list of all transactions, including credit usage for this account, displayed by month (default: current month). The transactions include:

* Bill by serverless: Credit usage from serverless GPU services
  * on-demand: Manual, function or server tasks charged at on-demand rates
  * spot: Tasks charged at spot pricing

### Dashboard

Provides credit usage overview. See [Dashboard](#dashboard) section for more details.

## Credit History

<figure><img src="/files/d6nnVEX13myV6elkPzvH" alt=""><figcaption><p>Credit History</p></figcaption></figure>

To view your credit history:

1. Go to "Settings"
2. Select the "Payment" section
3. Click on the "Credit" tab

You will see a complete history of all credits you have received.


# Workspace

Manage your workspace settings and configurations

## What is it?

A Workspace in Float16 App is a dedicated space where you can use various services separately from other spaces. It helps you manage resources across different projects or environments more effectively.

* **Environment Separation**: Ideal for managing different development environments (e.g., DEV, UAT, PROD) or multiple projects.
* **Workspace-Specific API Keys**: Each workspace has its own API keys
* **Isolated Analytics**: Analytics in each workspace show only transactions that occur within that workspace
* **Resource Management**: Helps in clearly organizing resources between projects or environments
* **Security Management**: Manage access permissions for each project or environment (Coming Soon)&#x20;

<figure><img src="/files/0mYh3VeybwGPX1Hoa7mX" alt=""><figcaption><p>All Workspace</p></figcaption></figure>

{% hint style="info" %}
Service quotas are account-based, not workspace-based. This means *all workspaces* within an account *share the same quota*.
{% endhint %}

## Manage Workspace

### Default Workspace

Upon your first login, a default workspace is automatically created for you. The default workspace name is derived from your email address. Example:&#x20;

* If your email is <john.doe@float16.cloud>, your default workspace name will be "john.doe".

You can change the name of your default workspace after creation. For instructions on how to change the workspace name, refer to the "[Change Workspace Configuration](#change-workspace-configuration)" section.

In Settings -> All Workspace, the default workspace will always be marked as "default".

{% hint style="warning" %}
The default workspace is a permanent workspace. *It cannot be deleted*.
{% endhint %}

### Switching Workspace

<figure><img src="/files/clB8U2ERFgls8sJTzWs0" alt=""><figcaption><p>Switching Workspace</p></figcaption></figure>

To switch from one workspace to another:

1. Click on the current workspace name located in the navigation bar, a list of all workspaces in your account will be displayed.
2. Click on the name of the workspace you wish to access.
3. The system will switch to your selected workspace and display its dashboard.

### Adding a Workspace

<figure><img src="/files/XfryplHwuus5ThKmI9yU" alt=""><figcaption><p>Adding a Workspace</p></figcaption></figure>

There are two ways to add a workspace:

1. Add a Workspace via Settings
   * Access "Settings" by clicking the app logo, your account<img src="/files/pz7AqJU7SjUmXJ2ribdo" alt="" data-size="line">in the bottom-left corner, or your profile in the top-right corner.
   * Navigate to "All Workspace"
   * Click "New Workspace" and enter a name for your workspace
   * Click "Create". Upon successful creation, you'll be automatically directed to your new workspace.
2. Add a Workspace using the Shortcut
   * Click on the workspace name located in the navigation bar. This will display a list of all workspaces in your account.
   * Click "Create Workspace" and enter a name for your new workspace
   * Click "Create". Upon successful creation, you'll be automatically directed to your new workspace.

{% hint style="warning" %}
For Serverless GPU, each account has a single token. When you add a new workspace, you will share the same token and project with all other workspaces in your account.
{% endhint %}

### Change Workspace Configuration

<figure><img src="/files/ePvAxk8MT4EIfuPvh7xC" alt=""><figcaption><p>Default Workspace's Setting</p></figcaption></figure>

Currently, the only available configuration change is modifying the workspace name.

#### Change Workspace's name

Follow these instructions to change a workspace's name:

1. Click on<img src="/files/q2739MvjGQl7so5KQL08" alt="" data-size="line">"Workspace Settings" located at the bottom-left corner of the page.
2. The "General" section will be displayed, showing the current workspace name and workspace ID.
3. Enter the desired new name in the "Workspace Name" field.
4. Click the "Save" button. A notification stating "Workspace Updated" will appear, confirming that the name has been successfully changed.

### Deleting Workspace

<figure><img src="/files/gytiLsZs4D7gGXAGbneX" alt=""><figcaption><p>Workspace Setting</p></figcaption></figure>

{% hint style="warning" %}

* The default workspace cannot be deleted.
* Once a workspace is deleted, this action cannot be undone.
  {% endhint %}

To delete a workspace, follow these steps:

1. Click on<img src="/files/q2739MvjGQl7so5KQL08" alt="" data-size="line">"Workspace Settings" located at the bottom-left corner of the page.
2. The "General" section will be displayed, locate the "Delete Workspace" button.
3. Click the "Delete Workspace" button, a confirmation dialog will appear.
4. Click "Delete" again in the confirmation dialog to proceed with the deletion.
5. A notification will appear stating "Workspace Deleted", you will be automatically redirected to the All Workspaces page.

## Theme

We offer Light and Dark themes to suit your preferences. Your chosen theme will be remembered and maintained even after logging out.

### How to Change Your Theme

1. Access "Settings" by:
   * Clicking the app logo or
   * Selecting your account<img src="/files/pz7AqJU7SjUmXJ2ribdo" alt="" data-size="line">in the bottom-left corner
   * Clicking your profile in the top-right corner, select "Settings"
2. Locate the theme options at the bottom-left of the Settings page.
3. Choose your preferred theme: <img src="/files/RXfqNQqoYOiFBm6kcEp0" alt="" data-size="line"> - Light  <img src="/files/tEm6tuU4fs1EUIKSlIh0" alt="" data-size="line">- Dark <img src="/files/fPsucSWfs1y0QV4xGni6" alt="" data-size="line"> - System (follows your device settings)

Your selected theme will be applied immediately and persist across sessions.


# Service Quota

Estimate your current quota usage

{% hint style="danger" %}
Service Quota is under maintenance
{% endhint %}

## What is Service Quota?

Service Quota in Float16.cloud refers to the limits on creating, using, or accessing our services for each account. These limits vary depending on the specific service.

## Your service quota

### LLM as a service

<figure><img src="/files/FnlHxfUhjVJTL6Ia3Pjt" alt=""><figcaption><p>Quota remaining</p></figcaption></figure>

We provide a service quota for our LLM as a service, specifically for API key creation:

* Maximum API keys per account: 20
* You can monitor overall usage on the Service Quota page
* Manage your API keys within each workspace's services section

To learn more about managing your API Key, please refer to our detailed guide: [Learn More About API Keys](/getting-started/llm-as-a-service/quick-start/set-the-credentials)

### One Click Deploy

We implement a service quota for our One-Click Deploy service based on GPU card usage:

* Maximum GPU card usage: 4 cards per account
* This limit applies regardless of GPU card type
* Track your overall usage on the Service Quota page

If you need to create a new instance but lack sufficient quota. You have to terminate an existing instance to reclaim the quota, quota is released immediately upon instance termination. Then proceed with creating your new instance.

For detailed instructions on creating and terminating instances, please refer to our comprehensive guide: [Learn More About One Click Deploy](/getting-started/one-click-deploy)

{% hint style="info" %}

#### Requesting More Quota

If you need additional quota, you can [contact us](<email:support@float16.cloud  >) to request an increase.
{% endhint %}

## Monitoring Your Quota

<figure><img src="/files/CKEFuvortlVJ9UBbcXoV" alt=""><figcaption><p>Service Quota Setting</p></figcaption></figure>

To check your current quota usage:

1. Go to the Settings section
2. Navigate to the Service Quota page
3. View your usage across all workspaces in your account

This page provides an overview of your quota utilization, helping you manage your resources effectively across your entire account.

You can also check your current quota usage and remaining quota directly from the service's credential menu.<br>


# LLM as a service

Seamless LLM Integration

{% hint style="danger" %}
This service is under maintenance.
{% endhint %}

## What is LLM as a service?

LLM as a Service is a large language model on-demand API service designed for users who want to utilize LLMs without deploying or managing the extensive resources they require. It's also suitable for those in the decision-making stage who wish to test models.

<figure><img src="/files/q6hOpobzhxwjyi50xQhv" alt=""><figcaption><p>LLM as a Service's Overview</p></figcaption></figure>

Our service provides instant API access, allowing users to:

1. Choose their desired LLM model
2. Obtain our API key
3. Implement the service immediately

Users can focus on developing their products without worrying about deployment complexities.

## Pricing

We charge based on the number of tokens used, with separate rates for prompt and completion tokens. Pricing varies between models.

For more detailed pricing information, please visit [this link](https://float16.cloud/product#pricing).

## Explore Use Case

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Q&#x26;A Bot</strong></td><td>Create your chatbot</td><td><a href="/pages/eqUbhq7l4KvY2g0O46KN">/pages/eqUbhq7l4KvY2g0O46KN</a></td></tr><tr><td><strong>Text-to-SQL</strong></td><td>Easily convert to SQL</td><td><a href="/pages/KaGTQdhZN3OghaGq9X63">/pages/KaGTQdhZN3OghaGq9X63</a></td></tr></tbody></table>


# Quick Start

LLM as a service quick start

## Setting up API Key

After accessing LLM as a service, you need to set up your API key. Learn how to set your API key [here](/getting-started/llm-as-a-service/quick-start/set-the-credentials).

## Quickly test API

To quickly try the API using cURL, use the following command:

```bash
curl -X POST https://api.float16.cloud/v1/chat/completions -d 

  '{
    "model": "seallm-7b-v3",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "สวัสดี"
      }
    ]
   }'

  -H "Content-Type: application/json" 
  -H "Authorization: Bearer <float16-api-key>"
```

Paste this in your terminal to see the response.

## Using the Chat API

Our API is compatible with OpenAI, allowing integration with your chat UI using OpenAI or LangChain libraries.

### OpenAI

1. Install the OpenAI package:

```bash
pip install openai
```

2. Use this Python code snippet (example using SeaLLM-7B-v2.5 model):

```python
import httpx
import openai

FLOAT16_BASE_URL = "https://api.float16.cloud/v1/"
FLOAT16_API_KEY = "<your API key>"

client = openai.OpenAI(
    api_key=FLOAT16_API_KEY,
    base_url=FLOAT16_BASE_URL,
)
client._base_url = httpx.URL(FLOAT16_BASE_URL)

# Streaming chat:
messages = [{"role": "system", "content": "You are truly awesome."}]

while True:
    content = input(f"User:")
    messages.append({"role": "user", "content": content})
    print(f"Assistant:", sep="", end="", flush=True)
    content = ""

    for chunk in client.chat.completions.create(
        messages=messages,
        model="seallm-7b-v3",
        stream=True,
    ):
        delta_content = chunk.choices[0].delta.content
        if delta_content:
            print(delta_content, sep="", end="", flush=True)
            content += delta_content
    
    messages.append({"role": "assistant", "content": content})
    print("\n")
```

For more information on the OpenAI library, visit the [OpenAI docs](https://platform.openai.com/docs/libraries/python-library).

### LangChain

To use Float16.cloud with the LangChain, follow these steps:

1. Install the LangChain package:

```bash
pip install langchain langchain_community
```

or

```bash
conda install langchain langchain_community -c conda-forge
```

2. Use this Python code snippet (example using SeaLLM-7B-v2.5 model):

```python
from langchain_community.chat_models import ChatOpenAI
from langchain.schema import HumanMessage

FLOAT16_BASE_URL = "https://api.float16.cloud/v1/"
FLOAT16_API_KEY = "<your API key>"

chat = ChatOpenAI(
    model="seallm-7b-v3",
    api_key=FLOAT16_API_KEY,
    base_url=FLOAT16_BASE_URL,
    streaming=True,
)

# Simple invocation:
print(chat.invoke([HumanMessage(content="Hello")]))

# Streaming invocation:
for chunk in chat.stream("Write me a blog about how to start to raise cats"):
    print(chunk.content, end="", flush=True)
```

For more information on the LangChain library, visit the [LangChain docs](https://python.langchain.com/v0.2/docs/integrations/chat/openai/).

{% hint style="info" %}
**For Further Assistance:**

If you need additional help, feel free to contact us at <support@float16.cloud>.
{% endhint %}


# Set the credentials

Set your API Key

## What is a Credential?

A credential is a unique secret key used by individual users. To utilize an API, you need this key to access or call the API. With proper permissions, the API will respond to your requests.

For our LLM as a service, an API key is required to access the API. Therefore, you must create your credential or API key before use.

## Managing API Keys

<figure><img src="/files/PBe7KsnQqre9lRNF5PVN" alt=""><figcaption><p>Credentials Management</p></figcaption></figure>

### Adding an API Key

To add a new API key, follow these steps:

1. In the service, navigate to the "Credentials" menu.
2. Click on "Generate New API Key".
3. Enter an API key name (optional) and click "Create API Key".
   * If you don't enter a name, it will default to "secret key".
4. Your API key is now successfully added.

### Deleting an API Key

To delete an existing API key:

1. In the service, navigate to the "Credentials" menu.
2. Locate the credential you wish to delete. Click<img src="/files/YlGyG9p3GlwLzqEqdn6X" alt="" data-size="line">at the end of the selected API key row.
3. Confirm the deletion.&#x20;
4. A notification will appear stating "API Key Revoked", indicating successful deletion.

## Quota Management

Currently, the system provides 20 api key quotas per account (across all workspaces). You can check your quota usage in the following ways:

* On the Credentials page
* When creating a new API key
* In the service quota settings

For more information about quota management, please refer to [Service Quota](/getting-started/account/service-quota).


# Supported Model

Float16.cloud Model Overview

Float16.cloud currently supports the following inference language models:

## SeaLLMs

* A Large Language Model (LLM) specialized for South East Asian (SEA) languages
* Supports 10 languages: English, Chinese, Vietnamese 🇻🇳, Indonesian 🇮🇩, Thai 🇹🇭, Malay 🇲🇾, Khmer 🇰🇭, Lao 🇱🇦, Tagalog 🇵🇭, Burmese 🇲🇲

### SeaLLM-7B-v2.5

API call model name: `seallm-7b-v2.5`

### SeaLLM-7B-v3

API call model name: `seallm-7b-v3`

{% hint style="info" %}
For more detailed information about models and their capabilities, please visit: [Hugging Face](https://huggingface.co/SeaLLMs), [Blog](https://blog.float16.cloud/thai-rag-with-llamaindex-weaviate-seallm/)
{% endhint %}

## SQLCoder

### SQLCoder-7B-v2

API call model name: `sqlcoder-7b-v2`

* A specialized model for Text-to-SQL conversion, offering improved speed and efficiency with a smaller model size
* Our interactive playground is available through [Quick Play](https://float16.cloud/texttosql), allowing you to experiment with the model's capabilities

{% hint style="info" %}
For more detailed information about models and their capabilities, please visit: [Hugging Face](https://huggingface.co/defog/sqlcoder-7b-2), [Blog](https://blog.float16.cloud/sqlcoder-7b-2/)
{% endhint %}

## Upcoming Models

* Meta-Llama-3.1-405B (Experiment)


# Limitation

API Limitation

<table><thead><tr><th>Limitation</th><th>Description</th><th data-hidden></th></tr></thead><tbody><tr><td>Request Rate Limit</td><td>120 Requests/Minute</td><td></td></tr><tr><td>Maximum Prompt Tokens/Request</td><td><ul><li>SeaLLM-7B-v2.5: 8k Tokens</li><li>SeaLLM-7B-v3: 32k Tokens</li><li>SQLCoder-7B-v2: 4k Tokens</li></ul></td><td></td></tr><tr><td>Maximum Completion Tokens/Request</td><td><ul><li>SeaLLM-7B-v2.5: 4k Tokens</li><li>SeaLLM-7B-v3: 4k Tokens</li><li>SQLCoder-7B-v2: 4k Tokens</li></ul></td><td></td></tr></tbody></table>


# API Reference

API Reference

## Chat Completions

<mark style="color:green;">`POST`</mark> `https://api.float16.cloud/v1/chat/completions`

#### Headers

| Name                                            | Type   | Description                                                  |
| ----------------------------------------------- | ------ | ------------------------------------------------------------ |
| authorization<mark style="color:red;">\*</mark> | string | Examples: `float16-api-123e4567-e89b-12d3-a456-426655440000` |

#### Request Body

| Name                                      | Type            | Description                                                                                                                                    |
| ----------------------------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| model<mark style="color:red;">\*</mark>   | string          | Enum (`"SeaLLM-7B-v2.5"`, `"SeaLLM-7B-v3"`, `"SQLCoder-7B-v2"`)                                                                                |
| message<mark style="color:red;">\*</mark> | array of object | <p>role (required) : Enum (<code>"system"</code>, <code>"user"</code>, <code>"assistant"</code>) </p><p></p><p>content (required) : String</p> |
| stream                                    | boolean         | default : false                                                                                                                                |
| max\_tokens                               | integer or null | Max Tokens (integer) or Max Tokens (null)                                                                                                      |

{% tabs %}
{% tab title="200: OK " %}

```json
{
  "id": "string",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "string",
        "role": "assistant",
      },
      "text": "string"
    }
  ],
  "created": 0,
  "model": "string",
  "object": "chat.completion",
  "usage": {
    "completion_tokens": 0,
    "prompt_tokens": 0,
    "total_tokens": 0
  }
}
```

{% endtab %}

{% tab title="422: Unprocessable Entity " %}

```json
{
  "detail": [
    {
      "loc": [
        "string"
      ],
      "msg": "string",
      "type": "string"
    }
  ]
}
```

{% endtab %}
{% endtabs %}


# One Click Deploy

{% hint style="danger" %}
This service is under maintenance.
{% endhint %}

## What is One Click Deploy?

One Click Deploy is a service that helps you **deploy LLMs without any configuration**, using just a Huggingface Model Repository. This service allows you to focus on model development while preventing chaos during the inference process.

The One Click Deploy service relies on **TensorRT-LLM** and the **Triton Inference Server**. (**NIM** is also available; please reach out to us for more information.)

## Why One Click Deploy with Float16 ?

Deploying LLMs requires consideration of inference speed, not just **batch size**, **maximum input length**, **number of tokens**, **quantization**, **Context caching**, and other factors.&#x20;

Float16 handles the complexity of configuring LLM deployment, ensuring you have the best experience with LLM serving.

### Main feature

* OpenAI Compatible
* Auto scheduler
* Long context support (128k)
* Quantization
* Context caching

## Pricing

One Click Deploy service charge based on **instance hours** like EC2. (whatever compute or not)

| GPU (number of card) | Region                   | Price per hours |
| -------------------- | ------------------------ | --------------- |
| L40sx1               | N. Virginia (us-east-1)  | $2.7            |
| L40sx1               | Oregon (us-west-2)       | $2.7            |
| L4x1                 | N. Virginia (us-east-1)  | $1.2            |
| L4x1                 | Oregon (us-west-2)       | $1.2            |
| A10x1                | Sydney (ap-southeast-2)  | $1.95           |
| A10x1                | Jakarta (ap-southeast-3) | $2.1            |
| A10x1                | Tokyo (ap-northeast-1)   | $2.2            |

L4x1 is mean the instance have NVIDIA GPU L4 1 card.

L4x4 is mean the instance have NVIDIA GPU L4 4 cards

## Use Case

#### Intensive workload

One Click Deploy provides a dedicated endpoint for you with **no rate limits** or additional costs.&#x20;

This endpoint is private and exclusive to your workload, ensuring it is not shared with others.

#### RAG

Leverage LLMs and vector search together to empower LLMs to access external knowledge or use your business's internal documents.

#### Multilingual

Proprietary solutions are not suitable for low-resource and specific language use cases. You can deploy models for specific languages, such as SeaLLM for South-East Asian languages, Typhoon and OpenThaiGPT for the Thai language.

#### Code co-pilot

An alternative to GitHub Co-Pilot, you can deploy your own co-pilot like CodeQwen1.5-7B-Chat and use it via Continue.dev to help with autocompletion and fill-in-the-middle coding tasks.


# Quick Start

One Click Deploy quick start

## Check all instances

<figure><img src="/files/Jl7T04Dra0T2zuLp1OSP" alt=""><figcaption><p>All instance page</p></figcaption></figure>

When you first access the One-Click Deploy service, you'll be presented with a table displaying all your instances, both active and inactive.

Every account is allocated a quota of 4 GPU cards, with no restrictions on the type of GPU. To check your current quota, simply click on the "Quota" button or navigate to the service quota settings.

{% hint style="info" %}
Learn more about service quota [here](/getting-started/account/service-quota)
{% endhint %}

## Add new instance

to start new instance:

<figure><img src="/files/lO1Imb9KFIKIcpfFJck2" alt=""><figcaption><p>Input model repository</p></figcaption></figure>

1. Paste Hugging Face model repository and token (if required), then click "Next"

<figure><img src="/files/rrukCW0xDNFwyt25cQDZ" alt=""><figcaption><p>Create Instance</p></figcaption></figure>

2. Review model name and input instance name.
3. Configure instance, select region and GPU type.
4. Review pricing and instance summary.
5. Click "Start Deploy", wait for "Start instance successfully" notification
6. Redirected to the instance's deployment section.

{% hint style="info" %}
We currently use basic optimization techniques. Learn more in the [technical](/getting-started/one-click-deploy/features) section. For model support limitations, check [here](/getting-started/one-click-deploy/limitation).
{% endhint %}

## Quickly test API

After successful deployment, test your model using:

### Instance Chat Playground

<figure><img src="/files/v3xwl50qxSMqr6Ccn40D" alt=""><figcaption><p>Chat Playground</p></figcaption></figure>

After successfully deploying your model, you can easily test its performance using the Instance Chat Playground. This convenient testing tool is readily accessible from your instance overview, allowing you to immediately interact with your deployed model.

* Default settings: temperature 0.5, max tokens 512
* Customize parameters, system prompt, and text message via GUI

### cURL

Or use following command:

```bash
curl -X POST http://api.float16.cloud/dedicate/JxlkeA5y2c/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <float16-api-key>" \
  -d '{
    "model": "<your model>",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant."
      },
      {
        "role": "user",
        "content": "สวัสดี"
      }
    ]
   }'
```

{% hint style="info" %}
Find copyable API formats (including OpenAI and LangChain) in the API tab.
{% endhint %}


# Instance Detail

Your instance detail

When you initiate deployment, the instance detail page becomes available. This page is divided into five sections, each providing crucial information about your deployed instance.

## Overview

<figure><img src="/files/Ic0rFV4JI750VVcPawKa" alt=""><figcaption><p>Overview</p></figcaption></figure>

The Overview section contains essential instance information:

* Model details (batch size, max input length, number of tokens)&#x20;
* Instance configuration (cloud provider, region, GPU type)&#x20;
* Endpoint and API key (visible after successful deployment)

{% hint style="info" %}
You can regenerate the API key for security purposes. [Learn more about API key management.](/getting-started/one-click-deploy/quick-start/re-generate-api-key)
{% endhint %}

{% hint style="info" %}
A playground is also provided for quick model testing. [See how to use the playground.](/getting-started/one-click-deploy/quick-start)
{% endhint %}

## Usage & Cost

<figure><img src="/files/hy4QlMPdrHJLH321C0IF" alt=""><figcaption><p>usage and cost log</p></figcaption></figure>

This section displays real-time usage and cost information:

* Each row represents usage and cost for a specific pricing period
* New rows are added when pricing changes

**Example:** If L4 GPU costs `1.00/hour` in September and increases to `1.20/hour` in October, you'll see separate rows for September and October usage.

{% hint style="info" %}
For a comprehensive view of all instance costs, visit the [payment settings](/getting-started/account/payment).
{% endhint %}

## Activity

The Activity section shows monthly instance activity. [Learn more about activity dashboard.](/getting-started/account/dashboard)

## Deployments

<figure><img src="/files/eBnUIpyiFgnqn3DNyA8S" alt=""><figcaption><p>deployment status</p></figcaption></figure>

This section displays the deployment status with the following possible states:

1. Initial: Checking model and limitations
2. Allocate: Allocating resources
3. Running: Deploy successful, ready to use
4. Terminated: Instance shut down (will still show 3 checked statuses)

{% hint style="warning" %}
If deployment fails, system will automatically terminate the instance
{% endhint %}

## Setting

Currently, the Setting section offers the option to terminate the instance. [Learn how to terminate an instance.](/getting-started/one-click-deploy/quick-start/terminate-instance)


# Re-generate API Key

Manage your API Key

<figure><img src="/files/eUPYAvgnAqXGkWwyDxBq" alt=""><figcaption><p>Re-generate your API Key</p></figcaption></figure>

Each instance is provided with one endpoint and one API key. While you cannot configure the endpoint name or API key, you have the option to re-generate the API key as needed.

To re-generate your API key:

1. Select the instance, navigate to the "Overview" section
2. Locate the "Re-generate" button next to the current API key
3. Click "Re-generate" to create a new API key

{% hint style="info" %}

* There is no limit to how many times you can re-generate your API key
* The new API key becomes active immediately upon re-generation
* The old API key becomes invalid instantly and cannot be used anymore
  {% endhint %}

{% hint style="warning" %}
Before re-generating your API key, ensure that you're ready to update any applications or services using the current key, as they will lose access once the new key is generated.
{% endhint %}


# Terminate Instance

Terminate your instance

<figure><img src="/files/ZvY9VTpLCEis35MibR8e" alt=""><figcaption><p>terminate instance</p></figcaption></figure>

To terminate an instance, follow these steps:

1. Navigate to the "All Instances" page
2. Select the instance you wish to terminate
3. Go to the "Settings" section
4. Locate and click the "Terminate Instance" button
5. A confirmation dialog will appear; click "Continue" to confirm termination
6. You will see a "Terminated Successfully" notification upon completion

{% hint style="warning" %}
Once terminated, an instance cannot be re-deployed or restarted. Please ensure you want to proceed before confirming termination.
{% endhint %}


# Features

All features are automatically enabled by default to improve overall performance.

### **OpenAI Compatible**

Supports OpenAI clients for chat and completion use cases, and Continue.dev for Co-Pilot use cases.

### **Batch size, Max input length, Number token**

Supports long contexts up to 1M context length and dynamic batching to improve utilization.

### **Quantization**

Reduces memory footprint to deploy LLMs to half the original model size.

Improves inference speed by 2 times.

### **Context caching**

Reduces redundant computation when requests have the same context.

Improves inference speed by 1.5 to 2 times on evaluation datasets and over 10 times in ideal scenarios.


# OpenAI Compatible

### Available Endpoint

`/{dedicate}/v1/chat/completions`

`/{dedicate}/v1/completions`

[Read the full details about request parameter](/getting-started/one-click-deploy/endpoint-specification)s

### Authentication

The endpoint provides an API key to access it.

### Rate limit

The endpoint does not have a rate limit, but it does have a limited batch size (similar to concurrency). The batch size is determined by GPU VRAM.

### Chat template

The endpoint has an option to automatically apply a chat template or use the prompt from the request. The chat template is determined by the `chat_template` attribute in the `tokenizer_config.json` file."


# Long context and Auto scheduler

One Click Deploy supports long contexts up to 1M context length.

### Batch size, (Context length) Max input length, Number token

<figure><img src="/files/Aq4S15ySflY43JIUcGGv" alt=""><figcaption></figcaption></figure>

Batch size, max input length, and number of tokens are interrelated and crucial when deploying LLMs by yourself.

#### A higher batch size increases the parallel LLM processes at the same time.

i.e. 1 batch size of 1 means processing one request at a time, and the next request will be processed only after the first one is completed.

#### A longer max input length increases the number of tokens in the prompt.

i.e. max input length of 4,096 means the request can use a maximum of 4,096 tokens per request.

#### The number of tokens determines the average maximum tokens per request.

i.e. with 16,384 tokens, an 8 batch size, and a 4,096 max input length,&#x20;

this endpoint should process a maximum of 8 requests simultaneously, provided the accumulated tokens of the 8 requests do not exceed 16,384 tokens.&#x20;

This means each request should have an average of no more than 2,048 tokens.

If the incoming request tokens exceed the number of tokens, the request will automatically be put into a queue and wait to be processed when enough tokens are available

### How does long context is work ?

#### 1. The maximum context length is read from `max_position_embeddings` in `config.json`.

If the model repository was trained with a long context, ensure the `max_position_embeddings` in `config.json` matches the max context used during training.

#### 2. Ensure the VRAM instance is sufficient.

Long contexts require more VRAM for inference.

The VRAM requirement scales linearly; doubling the context size requires double the VRAM.

#### 3. One Click Deploy automatically sets the maximum context length.

It estimates the maximum context length, batch size, and number of tokens using optimization techniques, model size, VRAM, and GPU compute compatibility.

### How does auto scheduler is work ?

The auto scheduler automatically triggers when certain scenarios are met:

#### 1. Exceeding batch size

i.e. If your instance can process 8 batch sizes, it means if you are processing 8 requests simultaneously and you have a new request, the new request will wait until a batch size is available.

{% hint style="info" %}
The new request does not wait for all batch sizes to complete; if a request in the batch size completes, the new request will automatically start processing.
{% endhint %}

#### 2. Exceeding number of token

The number of tokens is a hard cap to process parallel requests simultaneously to prevent OOM (Out-Of-Memory) errors.


# Quantization

Quantization is a technique to reduce the VRAM required for deploying LLMs.

Quantization has several parameters to consider because extreme quantization can lead to model collapse. Additionally, some quantization techniques, when used with the wrong inference engine, can slow down the model's inference speed.

One Click Deploy sets the default quantization to 8-bit weight. This technique minimizes the accuracy loss impact on the model and reduces the model size to half of the original.

For advanced quantization, we prioritize minimizing accuracy loss not only for a single language but also for multilingual performance.

### Evaluation Dataset

* M3Exam (Multiple-choice, Accuracy)


# Context caching

### What is context caching ?

Context caching is the ability to cache the KV (key-value) computations of the same context between requests.&#x20;

This feature helps speed up inference time by over 10 times when the request has the same context.&#x20;

The speedup is about 2 times when benchmarked with evaluation datasets like M3Exam.

### How does context caching is work ?&#x20;

Context caching is automatically triggered when requests have the same context within the same batch size.&#x20;

For example, if the endpoint has a batch size of 8 and each request has 1,024 tokens, and the requests share the same context for 900 tokens.

Instead of the system calculating 1,024 \* 8 = 8,192 tokens.

The system will calculate ((1,024 - 900) \* 8) + 1,024 = 2,016 tokens.&#x20;

This significantly reduces the compute required and improves the endpoint's latency.

### Use case

* RAG
* Few-shot prompting
* Code Co-pilot


# Limitation

### Supported Model

* Llama (1, 2, 3, 3.1)
* Mistral
* Qwen, Qwen2
* Gemma, Gemma2

### **Incoming model**

* RecurrentGemma
* Mamba

### **Regions**

Currently, we offer services across 5 AWS regions:

* North Virginia (us-east-1)
* Oregon (us-west-2)
* Tokyo (ap-northeast-1)
* Sydney (ap-southeast-2)
* Jakarta (ap-southeast-3)

### **GPU Types**

We currently support 3 types of GPU instances:

* NVIDIA L4
* NVIDIA L40s
* NVIDIA A10

Incoming GPU

* NVIDIA H100
* NVIDIA H200&#x20;

### Multi-GPU Support

At the moment, We do not support Multi-GPU Deployment.

Please note that our regional coverage and GPU options are subject to expansion in the future. We continuously strive to enhance our service offerings to meet evolving customer needs.


# Validated model

{% hint style="info" %}
One Click Deploy supports models based on specific model architectures. This means that if you fine-tune a model, we can also deploy it as long as it uses a supported architecture.
{% endhint %}

### Llama

* [meta-llama/Meta-Llama-3.1-8B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct)
* [scb10x/llama-3-typhoon-v1.5x-8b-instruct](https://huggingface.co/scb10x/llama-3-typhoon-v1.5-8b-instruct)
* [scb10x/llama-3-typhoon-v1.5-8b-instruct](https://huggingface.co/scb10x/llama-3-typhoon-v1.5-8b-instruct)

### **Mistral**

* [mistralai/Mistral-7B-Instruct-v0.2](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2)

### Qwen, Qwen2

* [SeaLLMs/SeaLLMs-v3-7B-Chat](https://huggingface.co/SeaLLMs/SeaLLMs-v3-7B-Chat)
* [SeaLLMs/SeaLLMs-v3-1.5B-Chat](https://huggingface.co/SeaLLMs/SeaLLMs-v3-1.5B-Chat)
* [Qwen/Qwen2-7B-Instruct](https://huggingface.co/Qwen/Qwen2-7B-Instruct)
* [Qwen/CodeQwen1.5-7B-Chat](https://huggingface.co/Qwen/CodeQwen1.5-7B-Chat) (FIM Supported)
* [Qwen/CodeQwen1.5-7B](https://huggingface.co/Qwen/CodeQwen1.5-7B) (FIM Supported)

### Gemma, Gemma2

* [SeaLLMs/SeaLLM-7B-v2.5](https://huggingface.co/SeaLLMs/SeaLLM-7B-v2.5)
* [google/gemma-2-2b-it](https://huggingface.co/google/gemma-2-2b-it)
* [google/gemma-2-9b-it](https://huggingface.co/google/gemma-2-9b-it)
* [google/gemma-2-27b-it](https://huggingface.co/google/gemma-2-27b-it)
* [google/codegemma-1.1-2b](https://huggingface.co/google/codegemma-1.1-2b) (FIM Supported)
* [google/codegemma-7b](https://huggingface.co/google/codegemma-7b) (FIM Supported)


# Endpoint Specification

### Endpoint

#### Available endpoint

`/{dedicate}/v1/chat/completions`

#### Method

Post

#### Header

Autherization : "Bearer {your-apikey}"

#### Request parameter

<table data-full-width="true"><thead><tr><th width="261.3333333333333">Parameter</th><th width="601">Description</th><th>Required</th></tr></thead><tbody><tr><td>messages</td><td><p>Messages must be in the same format as the OpenAI API format</p><p><br>[</p><p>{</p><p>"role" : "system", "user", "assistant"</p><p>},</p><p>{</p><p>"content" : "text"</p><p>}</p><p>]</p></td><td>Yes</td></tr><tr><td>model</td><td><p>The model should be from a Huggingface model repository.</p><p>i.e. SeaLLMs/SeaLLMs-v3-1.5B-Chat</p></td><td>Yes</td></tr><tr><td>stream</td><td>Boolean, By default is False</td><td>No</td></tr><tr><td>max_tokens</td><td>Int, By default is 1024</td><td>No</td></tr><tr><td>temperature</td><td>Float, By default is 0.7</td><td>No</td></tr><tr><td>repetition_penalty</td><td>Float, By default is 1.0</td><td>No</td></tr><tr><td>end_id</td><td>Int, By default uses eos_token_id from config.json of the model repository.</td><td>No</td></tr><tr><td>top_p</td><td>Float, By default is 0.7</td><td>No</td></tr><tr><td>top_k</td><td>Int, By default is 40</td><td>No</td></tr><tr><td>stop</td><td>Array, By default uses eos_token from config.json of the model repository.</td><td>No</td></tr><tr><td>random_sed</td><td>Int, By default is 2</td><td>No</td></tr></tbody></table>

`/{dedicate}/v1/completions`

#### Method

Post

#### Header

Autherization : "Bearer {your-apikey}"

#### Request parameter

<table data-full-width="true"><thead><tr><th width="261.3333333333333">Parameter</th><th width="601">Description</th><th>Required</th></tr></thead><tbody><tr><td>prompt</td><td><p>A prompt is the raw text and is not modified by applying a chat template.</p><p></p><p>A prompt must be used when you need to use a base model or a coding model. When using a coding model via the Continue.dev extension, the prompt will automatically be passed to the endpoint.</p></td><td>Yes</td></tr><tr><td>model</td><td>The model should be from a Huggingface model repository.<br>i.e. SeaLLMs/SeaLLMs-v3-1.5B-Chat</td><td>Yes</td></tr><tr><td>stream</td><td>Boolean, By default is False</td><td>No</td></tr><tr><td>max_tokens</td><td>Int, By default is 1024</td><td>No</td></tr><tr><td>temperature</td><td>Float, By default is 0.7</td><td>No</td></tr><tr><td>repetition_penalty</td><td>Float, By default is 1.0</td><td>No</td></tr><tr><td>end_id</td><td>Int, By default uses eos_token_id from config.json of the model repository.</td><td>No</td></tr><tr><td>top_p</td><td>Float, By default is 0.7</td><td>No</td></tr><tr><td>top_k</td><td>Int, By default is 40</td><td>No</td></tr><tr><td>stop</td><td>Array, By default uses eos_token from config.json of the model repository.</td><td>No</td></tr><tr><td>random_sed</td><td>Int, By default is 2</td><td>No</td></tr></tbody></table>


# Serverless GPU

## What is Serverless GPU ?

Serverless GPU is a service that provides GPU resources for short periods, such as 1 second, 5 seconds, or 30 seconds.&#x20;

Serverless GPU helps developers and data scientists rapidly create proofs of concept (POC) and move to production in the same environment.

## How to use ?

The Serverless GPU include comprehensive workloads like AI training, image and video analytics, LLM chatbots, or vector search.

Based on these use cases, our Serverless GPU supports development tasks such as code fixing, syntax correction, and POC.

Once your function is developed, you can directly deploy it to production with the same dependencies as in development mode.

#### Development Mode

* Development Mode is used to run code for non-server purposes. Use `float16 run <py>`

#### Spot Mode

* Spot Mode is sub-set of Development mode.&#x20;
* This mode design for cost effective while the price is lower than Developmet mode more than 10x.
* Can be interrupted at any time if system resources are needed for higher-priority tasks
* Use `float16 run <py> --spot`

<figure><img src="/files/GfzSTdIaj9s4sbZGNXVO" alt=""><figcaption><p>Float16 run with GPU</p></figcaption></figure>

#### Production Mode

* Production Mode is used to deploy code as a server. Use `float16 deploy <py>`&#x20;
* After deployment, Production Mode provides you with,
  * **Function Endpoint** : Container automatically stop after task completion. The cost are calculated during start and stop time.

<figure><img src="/files/yF9lQJQPg07BiG53z7HY" alt=""><figcaption><p>Float16 deploy with GPU</p></figcaption></figure>

### **How is it charged** ?&#x20;

Serverless GPU charges are based on 2 usage metrics.

### Compute time&#x20;

Compute time is the duration the serverless GPU is utilized, such as running tasks in development mode, requests in function mode, and starting server mode.&#x20;

The easiest way to inspect compute time is by checking the duration of each task via CLI or a web dashboard.

### Storage

Storage is measured in near real-time, with pricing based on GB/month

## Why Serverless GPU with Float16 ?

We ensure the performance of Serverless GPU by providing a high-performance end-to-end service.&#x20;

Our service is designed to minimize your effort without requiring code changes to migrate to using serverless GPU.

#### **No Cold start**

* Instant response under 100ms, Always-on instance.

#### No Vendor lock-in

* No code changes required, no decorators needed. Write native Python and enjoy coding.

#### Ultra-fast File storage

* High-performance storage supporting read speeds up to 10GB/s.

#### Realtime Debuging

* Real-time code updates, so you don't have to worry about caching issues.

## Pricing

\[May 2025] We now offer **Serverless GPU** access with pay-as-you-go pricing.

* **Instance Type:** NVIDIA H100
* **Pricing:**
  * **On-demand:** $0.006 per second (\~$21.6/hour)
  * **Spot:** $0.0012 per second (\~$4.32/hour)
* **Storage Fee:** $5.184 per GB/month

Feel free to try our service! We’d love your feedback via Discord or any Float16 community channel.

## Use Case

<table data-view="cards"><thead><tr><th></th><th></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td></tr><tr><td><strong>List all library</strong></td><td>Explore the complete set of libraries available in your container.</td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td></tr><tr><td><strong>Pre-load model</strong> </td><td>Accelerate model access using remote storage for improved performance.</td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td></tr><tr><td><strong>Server mode</strong></td><td>Create your own serverless server for scalable cloud-based applications.</td></tr></tbody></table>


# Quick Start

Serverless GPU quick start

Our Serverless GPU is accessed through a Command Line Interface (CLI). Upon visiting the serverless GPU service page in our app, you'll find an onboarding section that guides you through first-time use. We recommend following these suggestions for a smooth start.

For those eager to dive in, you can explore our use cases in the next section. However, please familiarize yourself with the following essential information first

## CLI Installation

Install the "Float16" CLI using one of the following methods, depending on your operating system:

### macOS

Use Homebrew to install Float16 CLI

```bash
brew install float16-cloud/float16/cli
```

### Windows or Linux

Install Float16 CLI globally using npm

```bash
npm install -g @float16/cli 
```

{% hint style="info" %}
Ensure you have the appropriate package manager (Homebrew for macOS or npm for Windows) installed on your system before proceeding with the installation. [Learn more](/journey/how-to-install-node)
{% endhint %}

## Verification

After installation, verify that Float16 CLI is correctly installed by running

```bash
float16 --version
```

This command should display the current version of Float16 CLI installed on your system.

## Get your token

You can now easily access your token directly from the Float16 web interface.

* Go to the Float16 Dashboard and click the **“Serverless GPU”** menu.
* Click the **“Token”** tab, copy the token shown there for use with the CLI.
* This token is required to authenticate any command using `float16`.

**Notes:**

* Tokens are for individual use only and must not be shared.
* If you're new to Float16, make sure to sign up and log in before accessing your token.

For any issues or if your token is missing, please contact our support team via Discord or the support page.

<figure><img src="/files/iGAIQpLqYqkUIlLLdDQa" alt=""><figcaption><p>Serverless authentication tokens</p></figcaption></figure>


# Mode

Serverless GPU Services

The Serverless GPU service offers two operational modes: Develop and Deploy. Each mode is designed for different use cases to help you efficiently utilize GPU resources.

### Development Mode

Development mode is optimized for one-time GPU tasks with immediate result delivery.\
Suitable for batch processing, short-term experiments, and analytical tasks.

**Command:**

```
float16 run <your_app> --name <name>
```

**Limitations:**

* Max execution time: 60 seconds per task
* Max concurrency: up to 8 tasks (not guaranteed; depends on system load)

**Pricing:**

* On-demand: $0.006 per second (\~$21.6/hour)

### Spot Mode (Within development mode)

Spot mode is designed for users who need longer tasks (> 60s), at lower cost.

**Note:** Spot tasks can be interrupted by on-demand tasks and resumed automatically when resources free up.

**Command:**

```sh
float16 run --spot --name <task_name> --budget <budget_value>
```

* Tasks may pause and resume based on resource availability
* No automatic task staging – users must handle resume logic

**Limitations:**

* No fixed execution time limit
* Max concurrency: up to 8 tasks (not guaranteed)

**Pricing:**

* Spot: $0.0012 per second (\~$4.32/hour)

### Production Mode

Production mode allows continuous deployment of GPU applications via API endpoints.\
Best suited for serving models and real-time inference.

**Command:**

```
float16 deploy <your_app> --project-id <project_id>
```

After deployment, you will receive:

* API endpoints for your app
* API key for authentication

**Endpoint Types:**

* **Function Endpoint**
  * Executes your task and shuts down the container after completion
  * Max execution time: 120 seconds per request

**Limitations:**

* Must be built with **FastAPI**
* One API key per deployment (regeneration supported)
* Max concurrency: up to 8 tasks (not guaranteed)

**Pricing:**

* 💵 On-demand: $0.006 per second (\~$21.6/hour)


# Task Status

Task Status Description

| Task Status                      | Mode                          | Status Explanation                                                                                                   |
| -------------------------------- | ----------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| Completed (On-demand)            | Development, Deployment       | Task completed successfully with no errors.                                                                          |
| Completed (Spot)                 | Spot                          | Spot task completed successfully with no errors.                                                                     |
| Completed (Failed)               | Development, Deployment       | Task execution failed.                                                                                               |
| Completed (Timeout)              | Development                   | Task did not complete within the allowed time limit (60 seconds).                                                    |
| Incomplete (Cancel)              | Development, Deployment       | Task was canceled by the user before execution (using command `float16 queue delete`).                               |
| Timeout                          | Deployment                    | Task did not respond within 120 seconds.                                                                             |
| Running                          | Development, Deployment, Spot | Task is currently running.                                                                                           |
| Pending                          | Development, Deployment, Spot | Task is in the queue, waiting to be processed.                                                                       |
| Waiting for resource             | Spot                          | Task was interrupted due to insufficient resources and is waiting for resources to become available before resuming. |
| Completed (Reached budget limit) | Spot                          | The task has reached the budget limit that you set.                                                                  |
| Stopped by user                  | Spot                          | The task was stopped by the user (using command `float16 task spot stop`).                                           |
| Failed (Invalid syntax)          | Spot                          | The task failed due to invalid syntax in the code.                                                                   |
| Failed (Unexpected indent)       | Spot                          | The task failed due to an unexpected indentation error in the code.                                                  |
| Failed (NameError)               | Spot                          | The task failed due to an undefined variable or function being used.                                                 |
| Failed (SyntaxError)             | Spot                          | The task failed due to a syntax error in the code.                                                                   |
| Failed (No module named)         | Spot                          | The task failed because a required module was not found or not installed.                                            |


# App Features

## Getting Started

### Prerequisites

* Obtain your Float16 authentication token
* Access the Float16 App for serverless GPU service

## Quickstart Guide

After obtaining your token, the Quickstart Guide provides a step-by-step introduction to develop mode usage and deployment mode procedures

<figure><img src="/files/JF9KnPMj8sMLRnPA1KTf" alt=""><figcaption><p>Quickstart Guide</p></figcaption></figure>

## Check all projects

<figure><img src="/files/XcG7adbztIGJdkoUhAx6" alt=""><figcaption></figcaption></figure>

The Projects section provides a comprehensive overview of all projects in your Float16 account.

* Displays all your projects.
  * Search by Project ID, and filter by status and creation date.
* **Click on any project** to view more detailed project information.

## Create new project

<figure><img src="/files/efUsI5E9ChBcGfVIx4oN" alt=""><figcaption><p>Create New Project</p></figcaption></figure>

Creating a new project in Float16 is simple and takes just a few steps:

* Click **“Create Project”** from the Projects page
* Select your **Instance Type** (currently available: `h100`)
* Input **"Project Name"** (Optional)
* Click **“Create Project”** to finalize the setup

Once created, your project will appear in the list under the **Projects** section. You can then deploy or run tasks using this project ID.

## Check all tasks

<figure><img src="/files/7jXzLVIOyBkQs2KX7Yo5" alt=""><figcaption></figcaption></figure>

The Tasks section offers a centralized view of all tasks across your entire Float16 account. You can efficiently manage tasks by filtering based on Task ID, Status, Type, and Date.

To explore task, click on individual tasks to access detailed information.

<figure><img src="/files/pDF9rctuDIcAx6sBaowh" alt=""><figcaption><p>Task Detail</p></figcaption></figure>


# Project Detail

Each active project in the Float16 web interface is organized into seven main sections:

## Overview

<figure><img src="/files/SmgMc0cnvvu2McukUG7w" alt=""><figcaption><p>Overview section</p></figcaption></figure>

The Overview section provides essential project details:

* Project Name
* Project ID
* Project instance type
* Project status
* Creation date
* Last update date
* Endpoint and endpoint status (if deployed)

If your project has a deployed endpoint:

* A 24-hour graph displays the number of requests and total usage time (in seconds)

If you have any active Spot tasks:

* They will be shown in this section as long as they are still running

## Deploy

<figure><img src="/files/IRud4uOCsLYGybDfCCaM" alt=""><figcaption><p>Deploy section </p></figcaption></figure>

The Deploy section displays comprehensive deployment-related information&#x20;

* Deployed date
* Endpoint (function and server)
* API key
* Request task history

### Endpoint Management Actions

#### Stop Endpoint

1. Click the "Stop" button while the endpoint is active
2. Confirm to stop the endpoint
3. Immediate endpoint termination

{% hint style="info" %}
Equivalent to CLI command: `float16 endpoint stop`
{% endhint %}

#### Start Endpoint

1. Click the "Start" button while the endpoint is inactive
2. Confirm to start the endpoint
3. Immediate endpoint activation

{% hint style="info" %}
Equivalent to CLI command: `float16 endpoint start`
{% endhint %}

#### Re-generate API Key

1. Click "Re-generate" button
2. Confirm regeneration
3. Automatically invalidates the old API key&#x20;

{% hint style="info" %}
Equivalent to CLI command: `float16 endpoint regenerate`
{% endhint %}

## Develop

<figure><img src="/files/N1s3HRnBO7PA7okQrzTt" alt=""><figcaption><p>Develop section</p></figcaption></figure>

The Develop section displays tasks executed using the `float16 run` command for the current project, including comprehensive task details.

## Storage

<figure><img src="/files/3Dc8MmbROw63ZqKxvBLX" alt=""><figcaption><p>Storage section</p></figcaption></figure>

The Storage section lists remote storage associated with the project.

{% hint style="info" %}

* Equivalent to `float16 storage ls` CLI command
  {% endhint %}

### Storage Management Actions

#### File Preview

You can also click on a selected file to preview its contents before taking any action.

<figure><img src="/files/DOs3DFKtIJuil13T0rRW" alt=""><figcaption><p>File Preview</p></figcaption></figure>

**Upload Files/Folders**

1. Click the "Upload" button.
2. Select a file or drag and drop files/folders into the dialog. You can upload multiple files at once.
3. Click "Upload". You may close the dialog while the upload is in progress.

{% hint style="info" %}
Equivalent to CLI command: `float16 storage upload -f <file>`
{% endhint %}

#### Download Files

1. Click ![](/files/lhSZx2JMSzv6SkvfQR3S) on the selected file
2. Click **"Download"**.
3. The file will be downloaded immediately.

{% hint style="info" %}
Equivalent to CLI command: `float16 storage download -f <file>`
{% endhint %}

#### Remove Files

1. Click ![](/files/lhSZx2JMSzv6SkvfQR3S) on the selected file or folder
2. Click **"Delete"**.
3. Confirm the deletion. The file will be removed immediately.

{% hint style="info" %}
Equivalent to CLI command: `float16 storage remove-on-remote -f <file>`
{% endhint %}

#### Copy Files

1. Click ![](/files/lhSZx2JMSzv6SkvfQR3S) on the selected file or folder
2. Click **"Copy to"**.
3. Select the destination project
4. (Optional) Enter the destination path.
   1. If you want to copy to the **root directory** of the selected project, you can leave this field empty.
5. Click **"Copy"** to confirm and start the copy process.

{% hint style="info" %}
Equivalent to CLI command: `float16 storage copy`
{% endhint %}

## Tasks

<figure><img src="/files/wIcC1aU5Z6FCMe3JCZW9" alt=""><figcaption><p>Project's tasks</p></figcaption></figure>

This section presents a comprehensive table of all tasks within the project, including both deployment and development mode tasks.

### Features

* Search and filter tasks by:
  * Task ID
  * Status
  * Type
  * Date
* Clickable task entries for detailed information

<figure><img src="/files/elcrZ9vn80l0CV73rSIJ" alt=""><figcaption><p>Task Details</p></figcaption></figure>

{% hint style="info" %}

* Refresh the website to update task information
* When a project is deleted, only Overview and Tasks sections remain accessible
  {% endhint %}

## Playground

The Playground section provides an interactive environment where you can develop, test, and deploy your applications directly in the browser—no local setup required.

Playground includes three tabs:

<figure><img src="/files/7GCkvS9SwTOIAWITJldX" alt=""><figcaption><p>Run playground</p></figcaption></figure>

#### Run

* Represents the Development Mode (same as `float16 run`)
* Write and edit a single `.py` file directly in the browser
* You can click **Run** to execute the script immediately on Serverless GPU
* Or use **Run Spot** for execute in spot mode
* You can also rename the script file (default: `quick-run.py`)

<figure><img src="/files/aDhBA59W2hXe2eFIAgNJ" alt=""><figcaption><p>Deploy playground</p></figcaption></figure>

#### Deploy

* Designed for deploying FastAPI apps (same as `float16 deploy`)
* Write your FastAPI app in a single `.py` file
* Click **Deploy** to launch the app and receive an API endpoint instantly

<figure><img src="/files/TNJZmTePBkWFn8Q5fWT4" alt=""><figcaption><p>Requirements playground</p></figcaption></figure>

#### Requirements

* Represents the `requirements.txt` file
* Add any additional Python packages your project needs
* Click **Install** to install the listed dependencies before running or deploying

{% hint style="info" %}
You can see your available credit and today's usage here. For full details and conditions, check [Payment](/getting-started/account/payment#credit-usage).
{% endhint %}

## Activity

<figure><img src="/files/N8LnQBaYQkrXuSAVL7am" alt=""><figcaption><p>Activity Dashboard</p></figcaption></figure>

The **Activity** page provides a visual summary of your GPU usage over time. For more information, please refer to the [Activity Dashboard](/getting-started/account/dashboard#serverless-gpu) section.

### Setting

<figure><img src="/files/0UhP5xDX6WhoBWFbeRWT" alt=""><figcaption><p>Setting</p></figcaption></figure>

This section allows you to manage basic project configurations.

* **Rename Project**: To rename your project, enter a new name in the input field and click the **Rename** button.
* **Delete Project**: To permanently delete the project, click the **Delete Project** button. *Please note that this action cannot be undone.*


# File storage

## Concept

<figure><img src="/files/7Yi6vyd7SnNXeyZdXE7z" alt=""><figcaption></figcaption></figure>

Each project have individual storage it **didn't share** between your projects.

{% hint style="info" %}
The file storage absolute path is `/apps/`

workspace : `/apps/workspace/`

background : `/apps/background/`

api-service : `/apps/api-service/`
{% endhint %}

`background` will contain you spot job and separate by they task\_id.

`workspace` will contain you run job with latest code only.

`api-service` will contain you deploy job with latest code only.

***

### Bandwidth cost

At the moment, It free.

***

### File features

<figure><img src="/files/6eCf1JFXtNjYTrOGCKZd" alt=""><figcaption></figcaption></figure>

#### Preview file

<figure><img src="/files/gtWV7JrcrE3HNxomR9UX" alt=""><figcaption><p>Quick preview when file less than 1 mb.</p></figcaption></figure>

When the file is less than 1 MB, the system will automatically generate a preview for you.&#x20;

However, if the file is larger than 1 MB, you will need to manually continue with the preview.

#### Others

<figure><img src="/files/8CE2gMnatg4IVCZmoMa4" alt=""><figcaption></figcaption></figure>

* Each file can be downloaded via the GUI.
* "Copy to" is used to copy a file between your projects.

***

### Absolute path

The file storage absolute path is `/apps/`&#x20;

workspace : `/apps/workspace/`

background : `/apps/background/`

api-service : `/apps/api-service/`

***

### Share model weight

We have pre-downloaded AI model weights as shown in the table.&#x20;

(Permission is read-only)&#x20;

If you are looking for a new model, please let us know via [Discord](https://discord.gg/j2DVTMjr67).

#### LLM

| Model                 | Absolute path                               | Option                                                            |
| --------------------- | ------------------------------------------- | ----------------------------------------------------------------- |
| Qwen3-4B              | /share\_weights/Qwen3-4B-GGUF/              | <p>Qwen3-4B-Q4\_K\_M.gguf</p><p>Qwen3-4B-Q8\_0.gguf</p>           |
| Qwen3-8B              | /share\_weights/Qwen3-8B-GGUF/              | <p>Qwen3-8B-Q4\_K\_M.gguf</p><p>Qwen3-8B-Q8\_0.gguf</p>           |
| Qwen3-14B             | /share\_weights/Qwen3-14B-GGUF/             | <p>Qwen3-14B-Q4\_K\_M.gguf</p><p>Qwen3-14B-Q8\_0.gguf</p>         |
| Qwen3-32B             | /share\_weights/Qwen3-32B-GGUF/             | <p>Qwen3-32B-Q4\_K\_M.gguf</p><p>Qwen3-32B-Q8\_0.gguf</p>         |
| Qwen3-30B-A3B         | /share\_weights/Qwen3-30B-A3B-GGUF/         | <p>Qwen3-30B-A3B-Q4\_K\_M.gguf</p><p>Qwen3-30B-A3B-Q8\_0.gguf</p> |
| Gemma3-12B            | /share\_weights/Gemma3-12B-GGUF/            | <p>gemma-3-12b-it-q4\_0.gguf</p><p>mmproj-model-f16-12B.gguf</p>  |
| Gemma3-27B            | /share\_weights/Gemma3-27B-GGUF/            | <p>gemma-3-27b-it-q4\_0.gguf</p><p>mmproj-model-f16-27B.gguf</p>  |
| typhoon2.1-gemma3-12b | /share\_weights/typhoon2.1-gemma3-12b-gguf/ | typhoon2.1-gemma3-12b-q4\_k\_m.gguf                               |

#### VLM

| Model          | Absolute path                                 | Option         |
| -------------- | --------------------------------------------- | -------------- |
| Qwen2.5-vl-7B  | /share\_weights/Qwen2.5-VL-7B-Instruct-exl2/  | \*.safetensors |
| Qwen2.5-vl-32B | /share\_weights/Qwen2.5-VL-32B-Instruct-exl2/ | \*.safetensors |
| UI-tars-1.5-7b | /share\_weights/UI-TARS-1.5-7B-exl2/          | \*.safetensors |

#### Embedding

| Model                | Absolute path                              | Option                          |
| -------------------- | ------------------------------------------ | ------------------------------- |
| Bge-m3               | /share\_weights/Bge-m3-GGUF/               | bge-m3-Q8\_0.gguf               |
| Qwen3-Embedding-0.6B | /share\_weights/Qwen3-Embedding-0.6B-GGUF/ | Qwen3-Embedding-0.6B-Q8\_0.gguf |
| Qwen3-Embedding-4B   | /share\_weights/Qwen3-Embedding-4B-GGUF/   | Qwen3-Embedding-4B-Q8\_0.gguf   |
| Qwen3-Embedding-8B   | /share\_weights/Qwen3-Embedding-8B-GGUF/   | Qwen3-Embedding-8B-Q8\_0.gguf   |


# Tutorials


# Hello World

Hello World with Float16 Serverless GPU

This tutorial will guide you through running your first "Hello World" program on Float16's serverless GPU platform. Follow these steps after installing the Float16 CLI.

## Step 1 : Login

First, authenticate with the Float16 CLI using your provided token

```bash
float16 login --token <YOUR_TOKEN>
```

## Step 2 : Initialize Your Project

Create a new project directory and navigate into it

```bash
mkdir firstProject
cd firstProject
```

Initialize the project with required configuration files

```bash
float16 init
```

This command creates a float16.conf file and an empty requirements.txt for managing dependencies.

## Step 3 : Create Your "Hello World" Script

<https://github.com/float16-cloud/examples/tree/main/official/run/helloworld>

Create a new Python file named "helloWorld.py" with a simple print statement:

{% code overflow="wrap" %}

```bash
echo 'print("Hello World! Welcome to Float16 Serverless GPU")' > helloworld.py
```

{% endcode %}

## Step 4 : Create project

Create a new project

```
float16 project create
```

{% hint style="info" %}
If you cannot create new project, [learn more](/getting-started/serverless-gpu/faq#why-cant-i-create-a-new-project)
{% endhint %}

## Step 5 : Start and Run

Start the remote container

```bash
float16 project start
```

Once successfully started, run your Python file on the remote instance

```bash
float16 run helloWorld.py
```

Wait for the execution to complete. Upon success, you'll see the following output

```bash
Hello World! Welcome to Float16 Serverless GPU
```

{% hint style="success" %}
Congratulations! You've successfully run your first program on Float16's serverless GPU platform.
{% endhint %}

## Explore More&#x20;

Ready to dive deeper? Check out these resources to expand your knowledge and capabilities with Float16.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# Install new library

Installing New Libraries in Your Float16 Remote Instance

This tutorial will guide you through adding new libraries to your Float16 remote instance. Since the remote instance starts as an empty container, you need to install libraries before running your code.

{% hint style="info" %}

* Float16 CLI installed and logged in
* If not logged in, refer to the [Hello World tutorial](/getting-started/serverless-gpu/tutorials/hello-world)
  {% endhint %}

This example scenario, we'll use a script 'cat.py' that imports the 'requests' library to fetch data from a cat fact API.

```python
import requests

cat_fact = requests.get('https://catfact.ninja/fact')
if cat_fact.status_code == 200:
    print("\nFun fact about cats:")
    print(cat_fact.json()['fact'])
```

While you can simply use `pip install requests` on your local machine, the process is different for the remote instance.

## Step 1 : Create requirements.txt

Create a requirements.txt file to specify dependencies.

```bash
float16 init
```

This command creates both 'float16.conf' and 'requirements.txt' files.

Open 'requirements.txt' and add the following

```
requests==2.31.0
```

This specifies the library name and the version you want to install.

## Step 2 : Start the Project

Start the container, which automatically installs packages listed in 'requirements.txt'

```bash
float16 project start
```

## Step 3 : Update Libraries (if needed)

After starting the project and successfully installing packages, you might find that the installed version doesn't match your local version. This could potentially affect your code. Here's how to update

{% hint style="info" %}
Check the current version: `pip show <PACKAGE> | grep Version` or `pip list`
{% endhint %}

To upgrade the version (e.g., from '2.31.0' to '2.32.3'), edit 'requirements.txt'

```
requests==2.32.3
```

Update package dependencies

```bash
float16 project install
```

## Step 4 : Run Your Script

Execute your script

```bash
float16 run cat.py
```

Example output (The response is randomly generated, so you may see a different cat fact.):

{% code overflow="wrap" %}

```bash
Fun fact about cats:
The average lifespan of an outdoor-only cat is about 3 to 5 years while an indoor-only cat can live 16 years or much longer.
```

{% endcode %}

{% hint style="success" %}
Congratulations! You've become proficient in installing new libraries in your Float16 remote instance.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# Prepare model weight

This tutorial will guide you through copying files from S3 or R2 to remote storage.

{% hint style="info" %}

* An active Float16 project
* AWS S3 storage credentials or Cloudflare R2
  {% endhint %}

## 1. Using Copy from S3 object compatibility

### Step 1 : Check Storage before copy

```
float16 storage ls
```

### Step 2 : Copy from S3 or R2

Use the following command to copy file from R2 to your remote storage.

{% code overflow="wrap" %}

```
float16 storage copy-to-remote \
    --path /workspace \
    --s3-uri https://40557a5d82224556015a5cxxx.r2.cloudflarestorage.com/BUCKET-A/weight-dir \ # This will download ALL files under "weight-dir" into "workspace"
    --s3-access-key 02e87e05dc0f3c27239d6d13bd6xxxxx \
    --s3-secret-key 2f44f34783f759613b24071afeac55b06286b5e401c9exxxxx \
    --s3_endpoint https://40557a5d82224556015a5cxxx.r2.cloudflarestorage.com
```

{% endcode %}

Replace the placeholders with your actual AWS S3 or R2 credentials:

* `<PATH>` : Your remote storage path
* `<YOUR_S3_URI_DESTINATION>`: Your R2 bucket original path or directory
* `<YOUR_S3_ACCESS_KEY>`: Your R2 access key
* `<YOUR_S3_SECRET_KEY>`: Your R2 secret key
* `<YOUR_S3_ENDPOINT>`: Your R2 endpoint #Optional if you using S3 don't paste this flag.

After running the command, check your specified S3 destination path. You should find the `output.txt` file successfully copied.

### Step 3 : Check Remote Storage

Check the files in your remote storage

```
float16 storage ls
```

***

{% hint style="success" %}
Congratulations! You've learned how to copy output files from Float16's remote storage to your AWS S3 storage.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# S3 Copy output from remote

get your output to S3

In the Float16 project, each project has its own remote storage that automatically stores Python files and output files generated during run or deploy commands. This tutorial will guide you through copying output files from remote storage to your personal AWS S3 storage.

{% hint style="info" %}

* An active Float16 project
* AWS S3 storage credentials
  {% endhint %}

## Step 1 : Prepare Your Script

Create a Python script (e.g., `test_output.py`) that generates an output file

```python
import os

class FileWriter:
    def __init__(self, folder_name="output_files"):
        self.output_dir = folder_name
        os.makedirs(self.output_dir, exist_ok=True)
    
    def get_file_path(self, filename):
        return os.path.join(self.output_dir, filename)
    
    def write_simple_text(self):
        filename = f"output.txt"
        with open(self.get_file_path(filename), 'w', encoding='utf-8') as file:
            file.write('Hi\n')
            file.write('This is Serverless GPU\n')
            file.write('from Float16')
        return filename

def main():
    writer = FileWriter("my_output_files")
    
    print("processing...")
    
    simple_file = writer.write_simple_text()
    print(f"- write {simple_file} successfully")
    
    print(f"\nOutput Path: {os.path.abspath(writer.output_dir)}")

if __name__ == '__main__':
    main()
```

## Step 2 : Run your script

First, start your project and then run the script

```
float16 project start
```

```
float16 run test_output.py
```

you will got successfully response from CLI.

## Step 3 : Check Remote Storage

Check the files in your remote storage

```
float16 storage ls
```

## Step 4 : Copy Output to S3

Use the following command to copy your output to AWS S3

{% code overflow="wrap" %}

```
float16 storage copy-output \
    --path my_output_files/output.txt \
    --s3-uri <YOUR_S3_URI_DESTINATION> \
    --s3-access-key <YOUR_S3_ACCESS_KEY> \
    --s3-secret-key <YOUR_S3_SECRET_KEY> \
    --aws-region <YOUR_S3_REGION>
```

{% endcode %}

Replace the placeholders with your actual AWS S3 credentials:

* `<YOUR_S3_URI_DESTINATION>`: Your S3 bucket destination path
* `<YOUR_S3_ACCESS_KEY>`: Your AWS access key
* `<YOUR_S3_SECRET_KEY>`: Your AWS secret key
* `<YOUR_S3_REGION>`: Your S3 bucket region

After running the command, check your specified S3 destination path. You should find the `output.txt` file successfully copied.

{% hint style="success" %}
Congratulations! You've learned how to copy output files from Float16's remote storage to your AWS S3 storage.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# R2 Copy output from remote

get your output to R2

In the Float16 project, each project has its own remote storage that automatically stores Python files and output files generated during run or deploy commands. This tutorial will guide you through copying output files from remote storage to your personal Cloudflare R2 storage.

{% hint style="info" %}

* An active Float16 project
* Cloudflare R2 storage credentials
  {% endhint %}

## Step 1 : Prepare Your Script

Create a Python script (e.g., `test_output.py`) that generates an output file

```python
import os

class FileWriter:
    def __init__(self, folder_name="output_files"):
        self.output_dir = folder_name
        os.makedirs(self.output_dir, exist_ok=True)
    
    def get_file_path(self, filename):
        return os.path.join(self.output_dir, filename)
    
    def write_simple_text(self):
        filename = f"output.txt"
        with open(self.get_file_path(filename), 'w', encoding='utf-8') as file:
            file.write('Hi\n')
            file.write('This is Serverless GPU\n')
            file.write('from Float16')
        return filename

def main():
    writer = FileWriter("my_output_files")
    
    print("processing...")
    
    simple_file = writer.write_simple_text()
    print(f"- write {simple_file} successfully")
    
    print(f"\nOutput Path: {os.path.abspath(writer.output_dir)}")

if __name__ == '__main__':
    main()
```

## Step 2 : Run your script

First, start your project and then run the script

```
float16 project start
```

```
float16 run test_output.py
```

you will got successfully response from CLI.

## Step 3 : Check Remote Storage

Check the files in your remote storage

```
float16 storage ls
```

## Step 4 : Copy Output to R2

Use the following command to copy your output to Cloudflare R2

{% code overflow="wrap" %}

```
float16 storage copy-output \
    --path /workspace/output.txt \
    --s3-uri https://40557a5d82224556015a5cxxx.r2.cloudflarestorage.com/BUCKET-A/ \
    --s3-access-key 02e87e05dc0f3c27239d6d13bd6xxxxx \
    --s3-secret-key 2f44f34783f759613b24071afeac55b06286b5e401c9exxxxx \
    --s3_endpoint https://40557a5d82224556015a5cxxx.r2.cloudflarestorage.com
```

{% endcode %}

Replace the placeholders with your actual Cloudflare R2 credentials:

* `<YOUR_S3_URI_DESTINATION>`: Your R2 bucket destination path
* `<YOUR_S3_ACCESS_KEY>`: Your R2 access key
* `<YOUR_S3_SECRET_KEY>`: Your R2 secret key
* `<YOUR_S3_ENDPOINT>`: Your R2 endpoint

After running the command, check your specified R2 destination path. You should find the `output.txt` file successfully copied.

{% hint style="success" %}
Congratulations! You've learned how to copy output files from Float16's remote storage to your Cloudflare R2 storage.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# Direct upload and download

Get Endpoint via Float16

This tutorial guides you through upload and download using Float16's storage.

{% hint style="info" %}

* Float16 CLI installed
* Logged into Float16 account
* VSCode or preferred text editor recommended
  {% endhint %}

## Step 1 : Create and start the project

```
float16 project create --instance h100
float16 project start
```

{% hint style="info" %}
If you didn't start the project, You can't use storage command before start the project.
{% endhint %}

## Step 2 : Prepare the script

<https://github.com/float16-cloud/examples/tree/main/official/spot/torch-train-and-infernce-mnist>

(download-mnist-datasets.py)

```python
import os
from torchvision import datasets, transforms

def download_mnist(data_path):
    if not os.path.exists(data_path):
        os.makedirs(data_path)
    
    transform = transforms.Compose([
        transforms.ToTensor(),
        transforms.Normalize((0.1307,), (0.3081,))
    ])

    # Download training data
    train_dataset = datasets.MNIST(root=data_path, train=True, download=True, transform=transform)
    
    # Download test data
    test_dataset = datasets.MNIST(root=data_path, train=False, download=True, transform=transform)

    print(f"MNIST dataset downloaded and saved to {data_path}")

if __name__ == "__main__":
    data_path = "../mnist-datasets"  # You can change this to your preferred location
    download_mnist(data_path)
```

## Step 3.1 : Upload via CLI

After downloaded. Use this command to upload datasets directory to remote path.

```python
float16 storage upload -f ./mnist-datasets -d datasets
```

{% hint style="info" %}

* The **storage upload** command use direct connect between your local machine direct to&#x20;
  {% endhint %}

## Step 3.2 : Upload via Website

<figure><img src="/files/Kkv9gXneHJfYYLS3GuvW" alt=""><figcaption></figcaption></figure>

## Step 4.1 : Download the file(s) via CLI

```
float16 storage download -f datasets -d ./local_datasets
```

## Step 4.2 : Download the file(s) via Website

<figure><img src="/files/eNvlhY8R8wyPedlLYics" alt=""><figcaption></figcaption></figure>

{% hint style="success" %}
Congratulations! You've successfully use your first server mode on Float16's serverless GPU platform.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# Server mode

Get Endpoint via Float16

This tutorial guides you through deploying a simple FastAPI "Hello World" application using Float16's deployment mode.

{% hint style="info" %}

* Float16 CLI installed
* Logged into Float16 account
* VSCode or preferred text editor recommended
  {% endhint %}

## Step 1 : Prepare Your Script

<https://github.com/float16-cloud/examples/tree/main/official/deploy/fastapi-helloworld>

(server.py)

```python
import os
import uvicorn
import asyncio
from fastapi import FastAPI
from fastapi.responses import JSONResponse
from utils import *
app = FastAPI()

@app.get("/hello")
async def read_root():
    return {"message": f"{say_hello() say_world()}"}

async def main():
    config = uvicorn.Config(
        app, host="0.0.0.0", port=int(os.environ["PORT"])
    )
    server = uvicorn.Server(config)
    await server.serve()
```

(utils.py)

```python
def say_hello():
    return "hello"
    
def say_world():
    return "world"
```

{% hint style="info" %}

* Save the script in a selected folder
* Navigate to the folder in your terminal
* Ensure the port is set to "port=int(os.environ\['PORT'])"
* Ensure the server is serve with "async def main"
  {% endhint %}

## Step 2 : Create project

```
float16 project create --instance h100
```

### Resulting Files

* `float16.conf`: Contains your project ID
* `requirements.txt`: Initially empty

{% hint style="info" %}
If you cannot create new project, [learn more](/getting-started/serverless-gpu/faq#why-cant-i-create-a-new-project)
{% endhint %}

## Step 3 : Deploy Script

```
float16 deploy server.py
```

After successful deployment, you'll receive:

* Function Endpoint
* Server Endpoint
* API Key

**Example:**

<pre><code><strong>Function Endpoint: http://api.float16.cloud/task/run/function/x7x2DFl8zU   
</strong>Server Endpoint: http://api.float16.cloud/task/run/server/x7x2DFl8zU       
API Key: float16-r-QoZU7uNlgDIFJ5IMrBtOCjuzVBlC

## curl
curl -X GET "{FUNCTION-URL}/hello" -H "Authorization: Bearer {FLOAT16-ENDPOINT-TOKEN}"

curl -X GET "http://api.float16.cloud/task/run/function/x7x2DFl8zU/hello" -H "Authorization: Bearer float16-r-QoZU7uNlgDIFJ5IMrBtOCjuzVBlC"
</code></pre>

## Step 4 : Endpoint Request

Use the provided endpoints with the API key (bearer token) to make requests.

**Endpoint Request Example:**

* Path: `/hello`
* Expected Response: `{"message": "Hello World!"}`

<figure><img src="/files/MzCesOKogd18QIGWSKIZ" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
To understand the differences between function and server modes, refer to [the dedicated section](/getting-started/serverless-gpu/quick-start/mode#deploy-mode).
{% endhint %}

{% hint style="success" %}
Congratulations! You've successfully use your first server mode on Float16's serverless GPU platform.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# LLM Dynamic Batching

Get Endpoint via Float16

This tutorial guides you through deploying dynamic batching with FastAPI application using Float16's deployment mode.

{% hint style="info" %}

* Float16 CLI installed
* Logged into Float16 account
* VSCode or preferred text editor recommended
  {% endhint %}

## What is dynamic batching ?

Deploying AI endpoints, also known as online serving, is crucial and challenging because it is difficult and complex to ensure proper VRAM mapping.&#x20;

During online serving, we have several techniques to enhance GPU utilization and increase throughput.&#x20;

One of the famous techniques is 'Dynamic batching.'&#x20;

Dynamic batching helps us maximize GPU utilization while solving the 'memory bound' issue.&#x20;

Dynamic batching allows you to pack incoming requests within a specific time frame, like 1 sec or 2 sec, into the same batch and perform inference simultaneously.

This helps us infer faster but might trade off with a slight increase in latency.

## Step 1 : Download and Upload the weight

We use [Typhoon2-8b](https://huggingface.co/scb10x/llama3.1-typhoon2-8b-instruct) (a fine-tuned version of Llama3.1-8b) to demonstrate.

```
huggingface-cli download scb10x/llama3.1-typhoon2-8b-instruct --local-dir ./typhoon2-8b/

float16 storage upload -f ./typhoon2-8b -d weight-llm
```

## Step 2 : Prepare Your Script

<https://github.com/float16-cloud/examples/tree/main/official/deploy/fastapi-dynamic-batching-typhoon2-8b>

(server.py)

```python
import os
import time
from typing import Optional
import uuid 
from fastapi import FastAPI
from fastapi.responses import JSONResponse
from pydantic import BaseModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import uvicorn
import asyncio

start_load = time.time()
model_name = "../weight-llm/typhoon-8b"
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name,padding_side="left")

app = FastAPI()

class ChatRequest(BaseModel):
    messages: str
    max_token : Optional[int] = 512
    

def process_llm(batch_data, batch_id):
    global model
    batch_tokenized = []
    for data in batch_data : 
        _text_formated = [{"role": "user", "content": data}]
        _text_tokenized = tokenizer.apply_chat_template(
            _text_formated,
            tokenize=False,
            add_generation_prompt=True
        )
        batch_tokenized.append(_text_tokenized)

    model_inputs = tokenizer(batch_tokenized, return_tensors="pt",padding=True,truncation=True).to(model.device)
    generated_ids = model.generate(
        **model_inputs,
        max_new_tokens=512,
        pad_token_id=tokenizer.eos_token_id
    )
    generated_ids = [
        output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
    ]

    result_list = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)
    result_with_id = dict(zip(batch_id,result_list))
    return result_with_id

class BatchProcessor:
    def __init__(self):
        self.batch = []
        self.batch_id = []
        self.lock = asyncio.Lock()
        self.event = asyncio.Event()
        
    async def add_to_batch(self, data, batch_id):
        async with self.lock:
            self.batch.append(data)
            self.batch_id.append(batch_id)

    async def process_batch(self):
        while True:
            await asyncio.sleep(1)  # Wait for 1 second
            async with self.lock:
                current_batch = self.batch.copy()
                current_batch_id = self.batch_id.copy()
                self.batch.clear()
                self.batch_id.clear()

            if current_batch:
                self.results = process_llm(current_batch,current_batch_id)
                self.event.set()
                self.event.clear()

    async def get_result(self, batch_id):
        return self.results[batch_id]

main_batch = BatchProcessor()

@app.post("/chat")
async def chat(text_request: ChatRequest):
    batch_id = uuid.uuid4()
    await main_batch.add_to_batch(text_request.messages, batch_id)
    await main_batch.event.wait()
    result_text = await main_batch.get_result(batch_id)
    return JSONResponse(content={"response": result_text})

async def main():
    asyncio.create_task(main_batch.process_batch())
    config = uvicorn.Config(
        app, host="0.0.0.0", port=int(os.environ["PORT"])
    )
    server = uvicorn.Server(config)
    await server.serve()
```

{% hint style="info" %}

* Ensure the port is set to "port=int(os.environ\['PORT'])"
* Ensure the server is serve with "async def main"
  {% endhint %}

## Step 3 : Deploy Script

```
float16 deploy server.py
```

After successful deployment, you'll receive:

* Function Endpoint
* Server Endpoint
* API Key

**Example:**

```
Function Endpoint: http://api.float16.cloud/task/run/function/x7x2DFl8zU   
Server Endpoint: http://api.float16.cloud/task/run/server/x7x2DFl8zU       
API Key: float16-r-QoZU7uNlgDIFJ5IMrBtOCjuzVBlC

## curl
curl -X POST "{FUNCTION-URL}/chat" -H "Authorization: Bearer {FLOAT16-ENDPOINT-TOKEN}"

curl -X POST "http://api.float16.cloud/task/run/server/x7x2DFl8zU/chat" -H "Authorization: Bearer float16-r-QoZU7uNlgDIFJ5IMrBtOCjuzVBlC" -d '{ "messages": "Hi !! Who are you ?" }' &
curl -X POST "http://api.float16.cloud/task/run/server/x7x2DFl8zU/chat" -H "Authorization: Bearer float16-r-QoZU7uNlgDIFJ5IMrBtOCjuzVBlC" -d '{ "messages": "How about you ?" }'
```

To pack requests, we need to use the server mode only.&#x20;

This is because server mode will start the endpoint and keep it alive for 30 seconds.&#x20;

(You will only be billed for 30 seconds)

It doesn't charge based on the number of requests during the active time.&#x20;

The server will handle and process the requests by itself. This will help you be more cost-effective.

{% hint style="success" %}
Congratulations! You've successfully use your first server mode on Float16's serverless GPU platform.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# Train and Inference MNIST

Get Endpoint via Float16

This tutorial guides you train and inference AI development using Float16's spot mode.

{% hint style="info" %}

* Float16 CLI installed
* Logged into Float16 account
* VSCode or preferred text editor recommended
  {% endhint %}

## Spot mode

Spot mode is cost effective for interruptable workload such as train AI model, offline inference, pre-processing and etc.

Spot mode is offer discount 80% when compare with run mode and server mode.

## Step 1 : Prepare Your Script

<https://github.com/float16-cloud/examples/tree/main/official/spot/torch-train-and-infernce-mnist>

(train.py)

```python
import torch
import torch.nn as nn
import torch.optim as optim
from torchvision import datasets, transforms
from torch.utils.data import DataLoader
import os

print(f"PyTorch version: {torch.__version__}")
# Define the neural network (same as before)
class Net(nn.Module):
    def __init__(self):
        super(Net, self).__init__()
        self.conv1 = nn.Conv2d(1, 32, 3, 1)
        self.conv2 = nn.Conv2d(32, 64, 3, 1)
        self.dropout1 = nn.Dropout2d(0.25)
        self.dropout2 = nn.Dropout2d(0.5)
        self.fc1 = nn.Linear(9216, 128)
        self.fc2 = nn.Linear(128, 10)

    def forward(self, x):
        x = self.conv1(x)
        x = nn.functional.relu(x)
        x = self.conv2(x)
        x = nn.functional.relu(x)
        x = nn.functional.max_pool2d(x, 2)
        x = self.dropout1(x)
        x = torch.flatten(x, 1)
        x = self.fc1(x)
        x = nn.functional.relu(x)
        x = self.dropout2(x)
        x = self.fc2(x)
        output = nn.functional.log_softmax(x, dim=1)
        return output

# Set device
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Device: {device}")
# Data loading
def load_data(data_path):
    print(f"load data : {data_path}")
    transform = transforms.Compose([
        transforms.ToTensor(),
        transforms.Normalize((0.1307,), (0.3081,))
    ])

    train_dataset = datasets.MNIST(root=data_path, train=True, download=False, transform=transform)
    test_dataset = datasets.MNIST(root=data_path, train=False, download=False, transform=transform)

    train_loader = DataLoader(train_dataset, batch_size=64, shuffle=True)
    test_loader = DataLoader(test_dataset, batch_size=1000, shuffle=False)

    return train_loader, test_loader

# Initialize the model, loss function, and optimizer
model = Net().to(device)
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)

# Training parameters
num_epochs = 10
save_interval = 2  # Save checkpoint every 2 epochs

# Checkpoint file path
checkpoint_path = 'mnist_checkpoint.pth'

# Function to save checkpoint
def save_checkpoint(epoch, model, optimizer):
    torch.save({
        'epoch': epoch,
        'model_state_dict': model.state_dict(),
        'optimizer_state_dict': optimizer.state_dict(),
    }, checkpoint_path)

# Function to load checkpoint
def load_checkpoint(model, optimizer):
    if os.path.exists(checkpoint_path):
        checkpoint = torch.load(checkpoint_path)
        model.load_state_dict(checkpoint['model_state_dict'])
        optimizer.load_state_dict(checkpoint['optimizer_state_dict'])
        start_epoch = checkpoint['epoch'] + 1
        print(f"Resuming training from epoch {start_epoch}")
        return start_epoch
    else:
        print("No checkpoint found. Starting training from scratch.")
        return 0

def train(model, train_loader, test_loader, num_epochs, save_interval):
    # Load checkpoint if it exists
    print(f"load checkpoint")
    start_epoch = load_checkpoint(model, optimizer)

    # Training loop
    for epoch in range(start_epoch, num_epochs):
        model.train()
        for batch_idx, (data, target) in enumerate(train_loader):
            data, target = data.to(device), target.to(device)
            optimizer.zero_grad()
            output = model(data)
            loss = criterion(output, target)
            loss.backward()
            optimizer.step()
            
            if batch_idx % 100 == 0:
                print(f'Epoch {epoch+1}/{num_epochs}, Batch {batch_idx}/{len(train_loader)}, Loss: {loss.item():.4f}')
        
        # Save checkpoint
        if (epoch + 1) % save_interval == 0:
            save_checkpoint(epoch, model, optimizer)
            print(f"Checkpoint saved at epoch {epoch+1}")

        # Evaluate on test set
        model.eval()
        test_loss = 0
        correct = 0
        with torch.no_grad():
            for data, target in test_loader:
                data, target = data.to(device), target.to(device)
                output = model(data)
                test_loss += criterion(output, target).item()
                pred = output.argmax(dim=1, keepdim=True)
                correct += pred.eq(target.view_as(pred)).sum().item()

        test_loss /= len(test_loader.dataset)
        accuracy = 100. * correct / len(test_loader.dataset)
        print(f'Epoch {epoch+1}/{num_epochs}, Test Loss: {test_loss:.4f}, Accuracy: {accuracy:.2f}%')

    # Save final model
    torch.save(model.state_dict(), 'mnist_model.pth')
    print("Training completed. Final model saved.")

data_path = "../datasets/mnist-datasets"  # Make sure this matches the path in download_mnist.py
train_loader, test_loader = load_data(data_path)
train(model, train_loader, test_loader, num_epochs, save_interval)
```

(inference.py)

```python
# inference_mnist.py

import torch
import torch.nn as nn
from torchvision import transforms
from PIL import Image
import numpy as np

# Define the same neural network architecture used for training
class Net(nn.Module):
    def __init__(self):
        super(Net, self).__init__()
        self.conv1 = nn.Conv2d(1, 32, 3, 1)
        self.conv2 = nn.Conv2d(32, 64, 3, 1)
        self.dropout1 = nn.Dropout2d(0.25)
        self.dropout2 = nn.Dropout2d(0.5)
        self.fc1 = nn.Linear(9216, 128)
        self.fc2 = nn.Linear(128, 10)

    def forward(self, x):
        x = self.conv1(x)
        x = nn.functional.relu(x)
        x = self.conv2(x)
        x = nn.functional.relu(x)
        x = nn.functional.max_pool2d(x, 2)
        x = self.dropout1(x)
        x = torch.flatten(x, 1)
        x = self.fc1(x)
        x = nn.functional.relu(x)
        x = self.dropout2(x)
        x = self.fc2(x)
        output = nn.functional.log_softmax(x, dim=1)
        return output

# Function to preprocess the image
def preprocess_image(image_path):
    transform = transforms.Compose([
        transforms.Resize((28, 28)),
        transforms.ToTensor(),
        transforms.Normalize((0.1307,), (0.3081,))
    ])
    image = Image.open(image_path).convert('L')  # Convert to grayscale
    image = transform(image).unsqueeze(0)  # Add batch dimension
    return image

# Function to load model and make prediction
def predict_digit(image_path, model_path):
    # Set device
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

    # Load the model
    model = Net().to(device)
    print(f"Loading model from {model_path}")
    model.load_state_dict(torch.load(model_path, map_location=device))
    model.eval()

    # Preprocess the image
    image = preprocess_image(image_path).to(device)

    # Make prediction
    with torch.no_grad():
        output = model(image)
        prediction = output.argmax(dim=1, keepdim=True)

    return prediction.item()

# Example usage
model_path = "mnist_model.pth"  # Path to your saved model
image_path = "../datasets/mnist-datasets/test"  # Path to the image you want to classify

# If you want to test with multiple images
test_images = ["../datasets/mnist-datasets/test/image1.jpg", "../datasets/mnist-datasets/test/image2.jpg", "../datasets/mnist-datasets/test/image3.jpg"]
result = []
for img in test_images:
    digit = predict_digit(img, model_path)
    result.append((img, digit))
print(f"Results: : {result}")
with open("mnist-results.txt", "w") as f:
    for img, digit in result:
        f.write(f"{img}: {digit}\n")
```

(download-mnist-datasets.py)

```python
import os
from torchvision import datasets, transforms

def download_mnist(data_path):
    if not os.path.exists(data_path):
        os.makedirs(data_path)
    
    transform = transforms.Compose([
        transforms.ToTensor(),
        transforms.Normalize((0.1307,), (0.3081,))
    ])

    # Download training data
    train_dataset = datasets.MNIST(root=data_path, train=True, download=True, transform=transform)
    
    # Download test data
    test_dataset = datasets.MNIST(root=data_path, train=False, download=True, transform=transform)

    print(f"MNIST dataset downloaded and saved to {data_path}")

if __name__ == "__main__":
    data_path = "../mnist-datasets"  # You can change this to your preferred location
    download_mnist(data_path)
```

## Step 2 : Download and Upload datasets

Download datasets into you local machine

```
python3 download-mnist-datasets.py
```

Upload MNIST datasets to the server

```
float16 storage upload -f ./mnist-datasets -d datasets
```

## Step 3 : Training the MNIST model

Start training the MNIST model with spot mode.

```
float16 run train.py --spot
```

### Resulting Files

* `mnist_checkpoint.pth`: MNIST model weight

## Step 4 : Inference the MNIST model

{% hint style="warning" %}
Make sure you have uploaded the test image to the server and changed the image path before running the inference.
{% endhint %}

```
float16 run inference.py
```

```
Results: : [
('../datasets/mnist-datasets/test/image1.jpg', 5), 
('../datasets/mnist-datasets/test/image2.jpg', 7), 
('../datasets/mnist-datasets/test/image3.jpg', 5)
]
```

{% hint style="success" %}
Congratulations! You've successfully use your first server mode on Float16's serverless GPU platform.
{% endhint %}

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# Etc.

Get Endpoint via Float16

<figure><img src="/files/Pfoudcy3E6GKtIo7K1zk" alt=""><figcaption></figcaption></figure>

For further examples,&#x20;

please visit our GitHub: <https://github.com/float16-cloud/examples>

## Explore More&#x20;

Learn how to use Float16 CLI for various use cases in our tutorials.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Hello World</strong></td><td>Launch your first serverless GPU function and kickstart your journey.</td><td><a href="/pages/YApNAqisqIbe17eFghB7">/pages/YApNAqisqIbe17eFghB7</a></td></tr><tr><td><strong>Install new library</strong></td><td>Enhance your toolkit by adding new libraries tailored to your project needs.</td><td><a href="/pages/yyygZPyF40EbYlvQ9c8d">/pages/yyygZPyF40EbYlvQ9c8d</a></td></tr><tr><td><strong>Copy output from remote</strong></td><td>Efficiently transfer computation results from remote to your local storage.</td><td><a href="/pages/BPAsR6EAOJexajZPitzW">/pages/BPAsR6EAOJexajZPitzW</a></td></tr><tr><td><strong>Deploy FastAPI Helloworld</strong></td><td>Quick start to deploy FastAPI without change the code.</td><td><a href="/pages/IapDkZzfM7mN9ePptZcb">/pages/IapDkZzfM7mN9ePptZcb</a></td></tr><tr><td><strong>Upload and Download via CLI and Website</strong></td><td>Direct upload and download file(s) to server.</td><td><a href="/pages/neh6VHbUO1k5sjbJcIiZ">/pages/neh6VHbUO1k5sjbJcIiZ</a></td></tr><tr><td><strong>More examples</strong></td><td>Open source from community and Float16 team.</td><td><a href="/pages/WofUCY294VqyFMOY62JX">/pages/WofUCY294VqyFMOY62JX</a></td></tr></tbody></table>

Happy coding with Float16 Serverless GPU!


# CLI References

## Authentication <a href="#authentication" id="authentication"></a>

### Login

Authenticate user to CLI via provided token.

```sh
float16 login --token <YOUR_TOKEN>
```

#### Options

* `--token <YOUR_TOKEN>` (required): Your authentication token.
  * If `--token` is provided, attempts to authenticate with the given token.
  * If `--token` is omitted, prompts the user to input the token.

{% hint style="warning" %}
All commands, except for `float16 init` , `float16 example list` ,`float16 example` ,require the user to be logged in before use.
{% endhint %}

### Token

Display the current token and configuration file path.

```sh
float16 token
```

### Logout

Log out the current user.

```bash
float16 logout
```

This command ends the current authenticated session.

## Project Management <a href="#project-management" id="project-management"></a>

### Init

Initializes a new project in the current directory. Creates `float16.conf` and `requirements.txt` files.

```bash
float16 init
```

### Example

creates a specific project from our predefined examples.

```bash
float16 example <EXAMPLE>
```

### Example list

displays a list of all available example projects.

```bash
float16 example list
```

{% hint style="info" %}
You can also view the list of example projects in our repository.
{% endhint %}

### Create

Creates a new project with a specific instance type to enable your project's computational environment.

**Requires**

* Instance type must be valid
* Project must not exist or must be deleted
* If `--instance` is omitted, interactive input for instance type will be prompted

```
float16 project create --instance <INSTANCE_TYPE> --name <PROJECT_NAME>
```

**Options**

* `--instance <INSTANCE_TYPE>` : (Required) Specifies instance type for container
* `--name <PROJECT_NAME>`  or  `-n <PROJECT_NAME>`: (Optional) Specifies name for project

### Start

Begins a new project session and automatically installs packages listed in requirements.txt (if present).&#x20;

#### Requires

* `float16.conf` file would exist.

```bash
float16 project start
```

### Install

Installs packages specified in the project's requirements.txt file.&#x20;

#### Requires

* `requirements.txt` file would exist.
* Project must be started before running this command.

<pre class="language-bash"><code class="lang-bash"><strong>float16 project install
</strong></code></pre>

### Delete <a href="#task-management" id="task-management"></a>

Removes project container, permanently deleting the project and associated resources.

**Requires**

* Instance type must be valid
* Project must not exist or be deletable
* If `--project-id` is omitted, reads project ID from `float16.conf`

```
float16 project delete
```

**Options**

* `--project-id` : Project ID to delete
  * Optional if project ID exists in float16.conf (will delete current active project)
  * Required if no project ID is found in float16.conf

### List

Displays list of all projects under the workspace, including deleted projects.

**Options**

* `-n, --limit <number>`: (Optional) Specifies the number of tasks to display.
* `-a, --all`: (Optional) Displays all tasks, overriding the default limit.

```
float16 project list
```

#### Output

```bash
Project ID | Instance Type | Status  | Created At          | Last Updated
aaaaaaaa   | aws           | Active  | 2023-09-24 10:00:00 | 2023-10-24 10:00:00
bbbbbbbb   | aws           | Deleted | 2023-09-24 11:30:00 | 2023-10-24 11:30:00
```

{% hint style="info" %}
Display 20 queued tasks if no option is specified
{% endhint %}

## Task Management <a href="#task-management" id="task-management"></a>

### Run

Executes a Python script on a remote instance.

#### Requires

* `<file>` must exist and be a .py file.
* `<name>` (if provided) must be alphanumeric and less than 64 characters.
* Project must be started before running this command.

```bash
float16 run <YOUR_FILE> --name <NAME>
```

#### Parameters

* `<YOUR_FILE>`: Path to the Python script to run.

#### Options

* `--name <NAME>`(Optional): Custom name for the task.
* `--spot` (Optional): Activate Spot mode.
* `--budget <budget_limit>`(Optional): \[Spot Mode] Config limit budget of spot task. (default 10 USD)

### Task

Retrieves details of a specific task.

```bash
float16 task <TASK_ID>
```

#### Parameters

* `<TASK_ID>` : task id must be exist&#x20;

#### Output

```bash
Task ID: TASK-XXXXX
Name: [task name]
Type: [Develop or Deploy]
Status: [status]
Created At: [timestamp]
Completed At: [timestamp or N/A]
Duration: [duration or N/A]
Project ID: [project id]
```

### Task list

Displays a list of all tasks.

#### Options

* `-n, --limit <number>`: Specify the number of queued tasks to display
* `-a, --all`: Display all queued tasks
* `--type` : Specify the type of tasks to display (e.g., manual, server, function).
* `--project-id` : Specify the project of tasks to display

```bash
float16 task list
```

#### Output

```bash
Task ID   | Name     | Type     | Status   | Created At          | Project ID
TASK-0001 | Task 1   | Develop  | Running  | 2023-09-24 10:00:00 | aaaaaaa
TASK-0002 | Task 2   | Deploy   | Completed| 2023-09-24 11:30:00 | aaaaa....bbbb
```

{% hint style="info" %}
Display 20 queued tasks if no option is specified
{% endhint %}

### Task log

Prints the log of a specific task.

#### Parameters

* `<TASK_ID>` : Existing task ID

```bash
float16 task log <TASK_ID>
```

#### Output

```
✅ Task log fetched successfully
Task Log:
====================
<YOUR_TASK_LOG>
====================
```

### Task stop

stop spot task

#### Parameters

* `<TASK_ID>` : Existing task ID

```bash
float16 task spot stop <TASK_ID>
```

### Task adjust

adjust spot task configuration

#### Parameters

* `<TASK_ID>` : Existing task ID

```bash
float16 task spot adjust <TASK_ID> --budget <budget_limit>
```

#### Options

* `--budget <budget_limit>`: Specify budget limit

## Queue Management <a href="#queue-management" id="queue-management"></a>

### Queue list

Displays a list of all queued tasks.

```bash
float16 queue list
```

#### Output

<pre class="language-bash"><code class="lang-bash">✅ Queued tasks fetched successfully
Queue ID  | Task ID | Created At | Position
<strong>0 | TASK-0003 | 2023-09-24 12:00:00 | QUEUE-001
</strong>1 | TASK-0004 | 2023-09-24 12:15:00 | QUEUE-002
</code></pre>

### Delete queue

Removes a specific task from the queue.

#### Parameters

* `<TASK_ID>` : Valid queue ID of the task to be removed

```bash
float16 queue delete <TASK_ID>
```

## Deployment <a href="#storage-management" id="storage-management"></a>

### Deploy

Deploys the specified application to the remote instance.

**Requires**

* Application must exist and be a .py file
* Project ID must exist

<pre><code><strong>float16 deploy &#x3C;YOUR_APP>
</strong></code></pre>

#### Parameters

* `<YOUR_APP>`: Path to the Python script to deploy.

#### Options

* `--project-id <project_id>`: Project ID for deployment
  * Optional if project ID exists in float16.conf
  * Required if no project ID is found in float16.conf

#### Output

```bash
Application deployed successfully
Endpoint: <Endpoint Func>
          <Endpoint Server>
API Key:  <Endpoint API Key>
```

### Endpoint

Lists all available endpoints for the current or specified project.

**Requires**

* Project ID must exist
* If `--project-id` is omitted, reads project ID from `float16.conf`

```
float16 endpoint
```

**Options**

* `--project-id` : Project ID to list endpoints
  * Optional if project ID exists in float16.conf (will list endpoints of current active project)
  * Required if no project ID is found in float16.conf

#### Output

```bash
Function Endpoint: <Endpoint Func>
Server Endpoint: <Endpoint Server>
API Key: <Endpoint API Key>
Project ID: <project-id>
Status: <Endpoint Status>
Last Deployed: YYYY-MM-DD HH:MM:SS
```

### Stop Endpoint

Stops the active endpoint for the current or specified project.

**Requires**

* Project ID must exist
* Endpoint status must be active
* If `--project-id` is omitted, reads project ID from `float16.conf`

```
float16 endpoint stop
```

**Options**

* `--project-id` : Project ID to stop endpoint
  * Optional if project ID exists in float16.conf (will stop current active project endpoint)
  * Required if no project ID is found in float16.conf

### Start Endpoint

Starts the inactive endpoint for the current or specified project.

**Requires**

* Project ID must exist.
* Endpoint status must inactive.
* If `--project-id` is omitted, read project ID from `float16.conf` instead.

```
float16 endpoint start
```

**Options**

* `--project-id` : Project ID must exist
  * Endpoint status must be inactive
  * If `--project-id` is omitted, reads project ID from `float16.conf`

### Re-generate API Key

Generates a new API key for the current or specified project.

**Requires**

* Project ID must exist
* If `--project-id` is omitted, reads project ID from `float16.conf`

```
float16 endpoint regenerate
```

**Options**

* `--project-id` : Project ID to re-generate API key
  * Optional if project ID exists in float16.conf (will re-generate current active project API key)
  * Required if no project ID is found in float16.conf

## Storage Management <a href="#storage-management" id="storage-management"></a>

### Storage list

Displays a list of files in the project.

#### Requires

* Project must be started before running this command.

```bash
float16 storage ls
```

#### Output

```bash
Filename       | Size    | Type | Last Modified
example.py     | 1.2 KB  | File | 2023-09-24 14:00:00
data/          | 4.0 MB  | Dir  | 2023-09-24 13:30:00
```

{% hint style="info" %}
The system can display up to 1,000 files.
{% endhint %}

### Copy output

Copies output files to the user's S3 bucket.

#### Requires

* Project must be started before running this command.
* All parameters must be valid.

{% code overflow="wrap" %}

```bash
float16 storage copy-output --path <PATH> --s3-uri <S3-URI> --s3-access-key <S3-ACCESS-KEY> --s3-secret-key <S3-SECRET-KEY> --aws-region <AWS-REGION>
```

{% endcode %}

#### Options

* `--path <PATH>` (required): Path to the file or directory to be copied
* `--s3-uri <S3-URI>` (required): S3 URI where the copied file or directory will be placed
* `--s3-access-key <S3-ACCESS-KEY>` (required): S3 access key for authentication
* `--s3-secret-key <S3-SECRET-KEY>` (required): S3 secret key for authentication
* `--aws-region <AWS-REGION>` (required): AWS region for the S3 bucket

### Copy file to remote instance

Copies files from user's S3 to the remote instance.

#### Requires

* Project must be started before running this command.
* All parameters must be valid.

{% code overflow="wrap" %}

```bash
float16 storage copy-to-remote --path <PATH> --s3-uri <S3-URI> --s3-access-key <S3-ACCESS-KEY> --s3-secret-key <S3-SECRET-KEY> --aws-region <AWS-REGION>
```

{% endcode %}

#### Options

* `--path <PATH>` (required): Path to the destination where the copied file or directory will be placed
* `--s3-uri <S3-URI>` (required): S3 URI where the files will be copied to
* `--s3-access-key <S3-ACCESS-KEY>` (required): S3 access key for authentication
* `--s3-secret-key <S3-SECRET-KEY>` (required): S3 secret key for authentication
* `--aws-region <AWS-REGION>` (required): AWS region for the S3 bucket

### Remove file on remote instance

Removes a file from the remote instance.

#### Requires

* Project must be started before running this command.
* `<FILE>` must be valid.

```bash
float16 storage remove-on-remote --destination <PATH>
```

#### Options

* `-f, --files <FILE>` (required): Path to the file or directory to be removed
* `-p, --project-id <PROJECT_ID>` (optional): Project ID of the file or directory to be removed

### Upload file to remote instance

Upload file or directory from local storage to remote instance.

#### Requires

* Project must be started before running this command.
* `<FILE>` must be valid.

```bash
float16 storage upload --files <FILE>
```

#### Options

* `-f, --files <FILE>` (Required): Path to the file or directory to be uploaded
* `-d, --destination <PATH>` (optional): Path to the file or directory to be placed
* `-p, --project-id <PROJECT_ID>` (optional): Project ID of the file or directory to be placed

### Download file to local storage

download file to local storage.

#### Requires

* Project must be started before running this command.
* `<FILE>` must be valid.

```bash
float16 storage download --files <FILE>
```

#### Options

* `-f, --files <FILE>` (Required): Path to the file or directory to be downloaded
* `-d, --destination <PATH>` (optional): Path to the file or directory to be placed
* `-p, --project-id <PROJECT_ID>` (optional): Project ID of the file or directory to be downloaded

### Copy file

Copy a file or folder from remote storage to a specified destination. This includes copying within the same project or across different projects.

#### Requires

* Project must be started before running this command.
* `<FILE>` must be valid.

#### Parameters

* `<ORIGIN_PROJECT_ID>`: Project ID of the source file or folder
* `<FILE_PATH>`: Path to the file or folder you want to copy
* `<DESTINATION_PROJECT_ID>`: Project ID of the destination
* `<DESTINATION_PATH>`: Destination path for the copied file or folder

{% code overflow="wrap" %}

```bash
float16 storage copy <ORIGIN_PROJECT_ID>:<FILE_PATH> <DESTINATION_PROJECT_ID>:<DESTINATION_PATH>
```

{% endcode %}

{% hint style="info" %}

* You can also use the `float16 storage cp` command as a shorthand.
* If a file or folder with the same name already exists at the destination, it will be replaced.
* If you're copying a file from the current project (as defined in `float16.conf`), you do **not** need to specify `<ORIGIN_PROJECT_ID>`. \
  **Example:** `float16 storage copy <FILE_PATH> <DESTINATION_PROJECT_ID>:<DESTINATION_PATH>`
* If the destination path is the root of the destination project, you may omit `<DESTINATION_PATH>`. \
  **Example:** f`loat16 storage copy <ORIGIN_PROJECT_ID>:<FILE_PATH> <DESTINATION_PROJECT_ID>:`
  {% endhint %}

## General <a href="#general" id="general"></a>

### Help

Displays general help information or help for a specific command.

For general help:

```bash
float16 --help
```

For command-specific help:

```bash
float16 <command> --help
```

#### Example

```bash
float16 run --help
```

This will display help information for the 'run' command

### Get version

Displays the current version of the Float16 CLI.

```bash
float16 -v
```

or

```bash
float16 --version
```


# FAQ

Frequently Ask Question about Serverless GPU

<details>

<summary>How can I find my project ID?</summary>

You can locate your project ID through several methods:

**Using CLI**:

* Run `float16 project list`
* Your current project ID will appear at the top of the table with "active" status

<img src="/files/IQr7kKn4RQw8t7xNMTvI" alt="" data-size="original">

**Checking float16.conf**:

* Look for the 'project\_id' parameter in your float16.conf file

![](/files/UoH9uHyLmNJJaPQIpx1o)

</details>

<details>

<summary>Why am I getting a "project ID not found" error?</summary>

This error occurs when the system cannot locate your project ID. Check the following:

1. Verify that you have a float16.conf file in your current directory
   * If missing, run `float16 init` to create it
2. Confirm that you have created a project
   * If not, use `float16 project create` to create one
   * The project ID will automatically update in your float16.conf file
3. If you created the project in another path or folder
   * Locate the `float16.conf` file in the directory where you last used the project
   * Manually copy the project ID from that configuration file
   * Apply the copied project ID to your current working directory's configuration
4. If you're still experiencing issues, contact our support team

</details>

<details>

<summary>Why am I getting "files not found" errors when running or deploying?</summary>

This error typically occurs when:

* The file is actually missing
* You're trying to access files in subfolders

You **must be** in the same directory as your target file when using run or deploy commands.

</details>

<details>

<summary>Why does my remote storage contain files I didn't copy?</summary>

The CLI automatically uploads certain files:

* All .py files in your directory during run or deploy operations
* Your requirements.txt file (if not empty) when you use `float16 project start`

</details>

<details>

<summary>When to use <code>float16 project start</code> ?</summary>

Imagine your project is like a car parked in a garage. The `float16 project start` command is essentially turning the key and warming up the engine. It's your go-to command when you want to breathe life into your project's environment.

Think of it as a wake-up call for your development container. After creating a new project or if you've been away for more than 24 hours, this command springs your environment back to action. It does two critical things: creates or reactivates your container and automatically installs all the libraries specified in your `requirements.txt`.

* New project? Run `float16 project start`
* Haven't coded in over a day? Use the command

**PS.** In deployment mode, you won't need this—the `float16 deploy` command handles container activation automatically.

</details>

<details>

<summary>Why am I not getting any response?</summary>

In the beta version, there is a computation time limit. If your computation exceeds this limit, you will receive a null response from the server.

</details>

<details>

<summary>Why am I getting "Endpoint is unavailable" response?</summary>

Your endpoint is currently inactive. When an endpoint is in an inactive state, it cannot process requests or return responses.

**How to Verify**:

* Use `float16 endpoint` command to view current status

**How to Resolve**:

```
float16 endpoint start
```

</details>

<details>

<summary>What is the difference between On-Demand and Spot tasks?</summary>

**On-Demand Task**

An on-demand task runs for a maximum of **60-120 seconds** without any interruptions. It supports three trigger types: **manual, function, and server**. These tasks can be executed in both **development mode** and **production mode**. Pricing is based on the **on-demand rate**.

**Spot Task**

A spot task runs in **Spot Mode**, which allows execution beyond **30 seconds** but can be **interrupted** if an **on-demand task** requires resources. Once the on-demand task is completed, the spot task **automatically resumes**. Pricing follows the **spot rate**, which is typically lower than the on-demand rate.

</details>


# Playground

Learn and play with us.

We are committed to supporting the developer community by simplifying your workflow, allowing you to focus on critical tasks while leaving complexity behind.&#x20;

To this end, we have developed no cost playground environments

**Float16 - Colab**

{% content-ref url="/pages/sMkwG5UVgSoBkz6qvNB5" %}
[Float16 - Colab](/getting-started/playground/float16-colab)
{% endcontent-ref %}

**FloatChat** (Archive)

A chat interface. Simply input your LLM API key to begin testing.

{% content-ref url="/pages/rFGROrIRevotePWeXp6t" %}
[FloatChat](/getting-started/playground/floatchat)
{% endcontent-ref %}

**FloatPrompt** (Archive)

An environment for exploring, testing, and sharing prompt examples.

{% content-ref url="/pages/kpGqwElCUgRSzhQ7ZvXQ" %}
[FloatPrompt](/getting-started/playground/floatprompt)
{% endcontent-ref %}


# FloatChat

Chat Playground for developers

<figure><img src="/files/3m5XCsvBJHZlYuSz0nl3" alt=""><figcaption><p>Chat Playground</p></figcaption></figure>

For developers seeking a platform to test API keys without implementing a user interface, we offer the Chat Playground. This versatile environment allows you to experiment with various language models in one convenient location.

* **Multiple Model Support**: OpenAI Models (GPT), Anthropic Models (Claude), All Float16 models and additional compatible models (must be compatible with OpenAI SDK)
* **Easy Setup**: Simply enter your API key to begin chatting with your chosen model.
* **Additional Functionality**: The playground includes several useful features such as Presets, Prompts and Tools.

## How to use

1. Navigate to the Chat Playground section.
2. Enter your API key for the desired model.
3. Select the model you wish to interact with.
4. Start chatting and testing your queries.

To access and start using the Chat Playground, please visit

{% embed url="<https://chat.float16.cloud>" %}

The Chat Playground environment and its features are provided free of charge. However, please note:

{% hint style="info" %}
While the playground itself is free to use, any *requests* made to the language models will incur *costs as per the usual pricing* of each model.
{% endhint %}

Feel free to utilize our playground for all your development and testing needs!


# FloatPrompt

Create Prompt, Run and Share with your colleague

<figure><img src="/files/48cw9BlZQ61Ba8zjYSJG" alt=""><figcaption><p>FloatPrompt</p></figcaption></figure>

For developers and non-developers alike, FloatPrompt offers a user-friendly environment to create, test, and refine prompts using few-shot techniques. This platform allows you to experiment with various models and share your results effortlessly.

* **Free Models:** Test your prompts using our provided models, including "SeaLLM-7b-v3" and "Eidy" (a specialized medical AI model).
* **OpenAI Integration:** Input your OpenAI API key to access models like GPT-4 and GPT-4 Mini.
* **Collaborative Sharing:** Easily share your prompts with colleagues via generated public URLs.
* **Customizable Parameters:** Adjust model settings for optimal results.

## How to use

1. Navigate to the FloatPrompt section.
2. Set the system prompt to define the AI's behavior.
3. Enter your user prompt or question.
4. Add multiple message pairs (assistant and user prompts) as needed.
5. Adjust model settings:
   * Select your preferred model (default: SeaLLM-7b-v3)
   * Set the temperature (default: 0.5)
   * Define max tokens (default: 512)
6. Click "Run" to see the model's response.
7. Use the "Share" feature to generate a public URL for your prompt.
8. Copy the link and share it with others.

To access and start using the FloatPrompt, please visit

{% embed url="<https://prompt.float16.cloud/prompt/new>" %}

Please Note:

{% hint style="info" %}
Prompts are *not automatically* *saved*. To preserve your work, generate a share link and save it for future modifications.
{% endhint %}


# Quantize by Float16

Coming Soon


# Float16 - Colab

<figure><img src="/files/u5cVUx6EOgy429oejbC6" alt=""><figcaption></figcaption></figure>

[Float16 - Colab](https://colab.float16.cloud/) is based on [Serverless GPU](https://app.float16.cloud/) and running with H100 GPU.

### Overview

Float16 - Colab is designed for running experiments and exploring GPU performance, such as inference AI models, fine-tuning AI models, vector search (NVIDIA RAPIDS), gene sequencing (NVIDIA Parabrick), and more.\
In addition to running code, Float16 - Colab supports collaboration between users.\
Float16 - Colab allows sharing workspaces with others and sharing both code and notes on GitHub.

### Concept

Float16 - Colab leverages serverless GPU resources and uses serverless GPU billing.\
Shared features include Projects, Tasks, and Code, but do not include Deployment or Storage.

| Features        | Float16 - Colab | Float16 - Serverless GPU |
| --------------- | --------------- | ------------------------ |
| Projects        | Yes             | Yes                      |
| Files           | Yes             | Yes                      |
| Tasks           | Yes             | Yes                      |
| Run Code (Spot) | Yes             | Yes                      |
| Storage         | No              | Yes                      |
| Deploy          | No              | Yes                      |

### Features

#### Projects

Projects are repositories for the user.\
Each project is separate from the others, including its own files and storage.\
We designed projects to function similarly to Git repositories

#### Files and Notes

Files are used to write code and specify which file to execute.\
Files also support importing other files within the same project, without special requirements.\
Notes serve as READMEs for each file and support Markdown formatting.

#### Run

The "Run" feature allows you to submit the current file for code execution.\
Submitted tasks are added to a queue and executed when resources are available.\
Credits or usage will be consumed based on the duration of the task.\
Only one task can run concurrently at a time.

#### Share

Project and task status can be shared across accounts.\
Share links allow switching between public and private projects as needed.

#### Import from Github

**Import via a repository.**\
Importing via a repository will bring every Python file (`.py`) in the first level of the repository into Colab.\
Example: <https://github.com/float16-cloud/examples/tree/main/official/run/helloworld>

**Import via repository JSON file.**

Importing via a repository JSON file provides greater granularity when importing repositories.\
This method does not import Python files directly but instead imports both Python files and their associated notes together.

\
This approach enhances the user experience when sharing projects.

```
{
  "fileList": [
    {
      "fileName": "app.py",
      "readme" : "app.md"
    }
  ]
}
```

Example here : <https://github.com/float16-cloud/examples/blob/main/official/run/helloworld/files.json>

### Pricing

Only the Serverless GPU service incurs costs; there are no additional charges for Float16 - Colab.

{% embed url="<https://float16.cloud/product/serverless>" %}
Serverless GPU Pricing
{% endembed %}

### Daily Credit

We run a campaign to allow users to access Float16 - Colab with a daily credit valued at $5 (reset at UTC+0, 23:59:59).\
After exhausting the credit, users will be charged based on their usage.


# Q\&A Bot (RAG)

#### Use Case Overview

This tutorial is based on the work done by Vulture Prime in creating a Q\&A Chatbot. The Chatbot in this case interacts with user-uploaded documents, answering questions based on the content.

<figure><img src="/files/fvEZVI1DMj6iORoqacOj" alt=""><figcaption></figcaption></figure>

### User Flow

<figure><img src="/files/VbZ6EGxMvC3oAEEkrQwj" alt=""><figcaption></figcaption></figure>

### Frontend Development

#### Implement the User Interface

In the frontend development phase of the Q\&A Chatbot project, we chose the following libraries and frameworks:

* **Next.js:** Installed using `create-next-app` for project setup.
* **Tailwind CSS:** Utilized for styling the user interface.
* **react-hook-form:** Employed for form management, reducing complexity in handling form states.
* **zod:** Used in conjunction with react-hook-form for type-safe input validation.

#### Connect Frontend with Backend

Understanding the backend API is crucial in this step. Key points include:

* Familiarity with API endpoints, data formats (parameters or body), and the sequence of steps.
* Handling CORS issues, ensuring the frontend and backend can communicate seamlessly.

API Interfaces obtained:

* **API LoadAndStore Interface:**

  * `GET {endpoint}/loadAndStore`
  * Response: 200 (OK) with information about the default URL, chunk size, chunk overlap, and collection name.

  **Request Body**

```json
{
  "url": "string", // default = https://lilianweng.github.io/posts/2023-06-23-agent
  "chunk_size": 1024, // default
  "chunk_overlap": 0, // default
  "collection_name": "string" // default = temp1
}
```

* **API Query Interfaces:**

  * `POST {endpoint}/queryWithOutRetrieval`
  * `POST {endpoint}/queryWithRetrieval`
  * Body: Contains the query and collection name.

  **Request Body**

```json
{
  "query": "string",
  "collection_name": "string"
}
```

#### Implement URL Input Validation

URL input validation is crucial to prevent incorrect data from being sent to the backend. This involved using `react-hook-form` and `zod` for frontend validation.

Example validation schema using `zod`:

```javascript
javascriptCopy codeexport const endpointScheme = z.object({
  endpoint: z.string().url().optional().or(z.literal('')),
});

export const collectionScheme = z.object({
  url: z.string({ required_error: 'Invalid url', invalid_type_error: 'Invalid url' }),
  // ...other fields
});

export const askScheme = z.object({
  query: z.string({ required_error: '', invalid_type_error: '' }),
  bot: z.string({}).optional(),
});
```

#### Implement Error Handling

Error handling for API calls was implemented to display meaningful messages on the website in case of failures. This ensures users can understand and troubleshoot issues easily.

#### Test, Optimize, and Deploy Frontend

After error handling, thorough testing of the frontend was conducted to ensure it met the project requirements. Optimization focused on improving code readability and user experience. The frontend was then deployed, involving building and addressing any errors before handing over deployment to the DevOps team.

[Link to Frontend Code](https://github.com/vultureprime/ai-web-interface/tree/main/next-rag-faqs)

### Backend Development

{% hint style="info" %}
How to create RAG with Langchain [Link](https://www.vultureprime.com/how-to/how-to-build-rag-with-langchain)
{% endhint %}

#### Set Up EC2 Server

The backend development started with setting up an EC2 server. Python was confirmed to be installed, and necessary libraries such as `langchain`, `chromadb`, `fastapi`, and `uvicorn` were installed.

```bash
python3 --version

pip install langchain
pip install chromadb
pip install fastapi
pip install "uvicorn[standard]"
```

#### Develop FastAPI

FastAPI was used to create the backend API. A simple initial API endpoint was created, and additional functionality like file upload was added.

Begin by creating a `main.py` file with a simple endpoint:

```python
from fastapi import FastAPI

app = FastAPI()

@app.get("/")
def root():
    return {"message": "Hello World"}
```

To add more functionality, you can implement methods like POST for file uploads.

```python
@app.post("/upload")
def update():
    # ...your code...
    return {"message": "Uploaded"}
```

Run FastAPI using:

```bash
uvicorn main:app
```

#### Integrate with RAG

Integration with RAG involved understanding the main steps: importing data and querying data. Functions for uploading data and querying data were developed and associated with corresponding API methods.

#### Chroma Vector Database

Chroma vector database was integrated to enhance data querying capabilities. The `chromadb` library was installed and configured as the vector database for Langchain.

#### Allow CORS

CORS (Cross-Origin Resource Sharing) was configured to enhance security for frontend integration. Add the following middleware in FastAPI:

```python
from fastapi.middleware.cors import CORSMiddleware

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)
```

#### Deploy Your Code

The backend code was deployed, and the use of tools like `screen` for detached execution was explained.

```bash
# Create a new screen session
screen -S name

# Navigate to the API folder
cd path/to/api

# Start FastAPI
uvicorn main:app

# Detach from the screen session
Ctrl+a d
```

Ensure that the security group of the EC2 instance allows traffic on the specified port (default is 8000).

#### Link FastAPI with API Gateway

AWS API Gateway was introduced to manage traffic and authentication for added security and features.

#### Test and Tuning

The backend was tested by querying data, and tuning involved adjusting parameters like data division and AI creativity to achieve desired results.

[Link to Backend Code](https://github.com/vultureprime/deploy-ai-model/tree/main/paperspace-example/openai-langchain-basic-RAG)


# Text-to-SQL

#### Use Case Overview

The Text-to-SQL Chatbot is designed for Business Analysts (BA) or Product Managers (PM) who need to query a PostgreSQL database without expertise in SQL. The frontend provides a user-friendly interface, allowing users to input queries in natural language and receive corresponding SQL queries and database results.

<figure><img src="/files/RgetcFo1wpfsbcbzois7" alt=""><figcaption></figcaption></figure>

### User Flow

1. **Connect to Database:**
   * Users paste the Backend (BE) endpoint connected to the PostgreSQL database.
   * After clicking connect, the website displays the database schema and data.
2. **Write Queries:**
   * Users type prompts or questions into the chat.
   * The website responds with the corresponding SQL queries, displaying them in the chat.

![User Flow](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/655851b2a4ac02b28eb71686_Text-to-sql-6.png)

### Frontend Development

#### Frontend Development with Next.js

The frontend is implemented using Next.js, with key libraries including `react-hook-form`, `@tanstack/react-query`, `Tailwind CSS`, and `zod` for schema validation.&#x20;

create project

```bash
npx create-next-app
```

```
yarn dev
```

Form by zod

```tsx
export const askScheme = z.object({
  endpoint: z.string().url().optional().or(z.literal('')),
  query: z.string(),
  bot: z.string({}).optional(),
})
```

react-hook-form

```tsx
const methods = useForm <IOpenAIForm>({
    resolver: zodResolver(askScheme),
    mode: 'onChange',
    shouldFocusError: true,
    defaultValues: {
      endpoint: endpointAPi,
    },
  })

  const {
    handleSubmit,
    setError,
    setValue,
    control,
    formState: { errors },
  } = methods
```

The integration involves connecting to the backend API using `axios` and `@tanstack/react-query`. Data, including column information and random data for the table, is fetched from the backend, and the UI is updated accordingly.

```tsx
const [
    { isLoading, data: column, refetch: refetchColum },
    { isLoading: isLoadAlldata, data: allData, refetch: refetchAllData },
  ] = useQueries({
    queries: [
      {
        queryKey: ['getInfo'],
        queryFn: () => axios.get('/getInfo').then((res) => res.data),
      },
      {
        queryKey: ['getAllData'],
        queryFn: () => axios.get('/getAllData').then((res) => res.data),
      },
    ],
  })
```

```tsx
const { mutateAsync: queryPromt } = useMutation({
    mutationFn: ({ query_str }: { query_str: string }) => {
      return axios.get(`/queryWithPrompt?query_str=${query_str}`)
    },
    onError: () => {
      setError('bot', {
        message: 'There was an error fetching the response.',
      })
    },
  })

  const { mutateAsync: randomData } = useMutation({
    mutationFn: () => {
      return axios.get(`/addRandomData`)
    },
  })
```

#### Deploying the Frontend

To deploy the frontend, configure environment variables such as the API endpoint in the `.env` file. The tutorial provides guidance on deploying using Next.js.

```
NEXT_PUBLIC_API = your endpoint
```

[Link to Frontend Code](https://github.com/vultureprime/ai-web-interface/tree/main/next-text-to-sql)

### Backend Development

#### Setting up AWS EC2 Instance

This section guides you through setting up an AWS EC2 instance for database and backend API installation. The recommended EC2 type is one with sufficient RAM, such as `m5.large`, `c5.large`, or `r5.large`, along with a minimum of 50GB storage. The chosen operating system is Ubuntu.

#### Installing and Configuring PostgreSQL on EC2

The tutorial suggests using Docker for easy installation of PostgreSQL. It provides step-by-step commands for installing Docker on the EC2 instance and running a PostgreSQL container.

```python
# Add Docker's official GPG key:
sudo apt-get update
sudo apt-get install ca-certificates curl gnupg
sudo install -m 0755 -d /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
sudo chmod a+r /etc/apt/keyrings/docker.gpg

# Add the repository to Apt sources:
echo \
  "deb [arch="$(dpkg --print-architecture)" signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu \
  "$(. /etc/os-release && echo "$VERSION_CODENAME")" stable" | \
  sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
sudo apt-get update
```

```bash
sudo apt-get install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
```

install Postgres by Docker

```bash
docker run --name llm-postgres -e POSTGRES_PASSWORD=mysecretpassword -p 5432:5432 -d postgres
```

#### Setting up FastAPI Backend

**Initial FastAPI**

FastAPI is chosen for creating the backend API. The tutorial introduces the basic structure of a FastAPI project and demonstrates enabling CORS for frontend communication.

We will use FastAPI to create an API. Llamaindex serves as the intermediary for handling data models and the LLM model, which is GPT-4. Therefore, it is necessary to create an API key for OpenAI before using it, which can be done at <https://platform.openai.com/>.

necessary python library for FastAPI

```bash
pip install fastapi
pip install "uvicorn[standard]"
```

create app.py and initial FastAPI

```python
from fastapi import FastAPI
from fastapi.encoders import jsonable_encoder
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware
app = FastAPI()

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)


@app.get("/helloworld")
async def helloworld():
    return {"message": "Hello World"}
```

**Initial PostgreSQL Connection**

The backend connects to PostgreSQL using SQLAlchemy and psycopg2. Environment variables are loaded using `python-dotenv`, and key parameters for connecting to the database are declared.

Install necessary libraries using pip

```bash
pip install sqlalchemy psycopg2-binary faker numpy
```

Declare parameters for creating a database connection, pulling data from environment variables

```python
import os
import dotenv

dotenv.load_dotenv()
HOST = os.environ['HOST']
DBPASSWORD = os.environ['DBPASSWORD']
DBUSER = os.environ['DBUSER']
DBNAME = os.environ['DBNAME']
```

Use the provided GitHub example to set up an API for preparing tables and data related to student information, including height and weight.

The GitHub example includes the following API endpoints:

* `/createTable`: Create the student table.
* `/getInfo`: Check information about the table.
* `/addRandomData`: Add sample data to the table.
* `/getAllData`: Retrieve all data from the table.

Optionally, the example includes an endpoint for removing the created table:

* `/removeTable`: Delete the table created.

**Integrating LlamaIndex and OpenAI**

LlamaIndex is used for data modeling, acting as an intermediary between the database and the LLM (GenerativeAI) model. The tutorial covers the installation of LlamaIndex and the setup of OpenAI API keys.

Install Llamaindex Library

```bash
pip install llama_index
```

Declare OpenAI API Key as Environment Variable

```
OPENAI_API_KEY=your_key
```

#### Deploying and Testing the Backend

The backend server is deployed, and testing is performed using various API endpoints. The tutorial shows how to create tables, add random data, and query with prompts to retrieve results from the database.

#### Deployment Steps:

1. Create a Screen Session

   ```bash
   screen -S name
   ```
2. Navigate to the API Folder

   ```bash
   cd path/to/api
   ```
3. Start FastAPI

   ```bash
   uvicorn app:app
   ```
4. Detach from Screen Session

   ```bash
   Ctrl+a d
   ```

#### Testing Endpoints:

1. **Test `/createTable` Endpoint:**
   * **Response:**

     ```json
     {
         "message": "complete"
     }
     ```
2. **Test `/addRandomData` Endpoint:**
   * **Response:**

     ```json
     [
         {
             "id": 279703,
             "name": "Christina",
             "lastname": "Santos",
             "height": 201.53,
             "weight": 68.07
         },
         // ... (other data entries)
     ]
     ```
3. **Test `/queryWithPrompt` Endpoint:**
   * **Prompt:**

     ```
     "Write SQL in PostgreSQL format. Get average height of students."
     ```
   * **Response:**

     ```json
     {
         "result": "The average height of students is approximately 176.56 cm.",
         "SQL Query": "SELECT AVG(height) FROM students;"
     }
     ```

#### Setting up AWS API Gateway

The final section demonstrates integrating the deployed API with AWS API Gateway for enhanced management capabilities, such as authentication. For more information: [Link](https://www.vultureprime.com/blogs/api-gateway)

[Link to Backend Code](https://github.com/vultureprime/ai-web-backend/tree/main/Text-to-sql-openai-postgresSQL)


# OpenAI with Rate Limit

#### Use Case Overview

we will refer to the basics demonstrated in the demo created by Vulture Prime. It is a website to generate API keys for users who want to connect to Vulture Prime's OpenAI. Users can generate API keys through the website automatically without the need for admin approval. Users can also choose the desired rate limit per day.

Note: In cases where there is a user dashboard, users need to log in before using the system.

![Image](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/65619604c7d1f8d7a3be05da_po_1.png)

### User Flow

To facilitate management, we have created a UI page that allows users to access the API key easily without contacting the admin. Users can obtain the OpenAI API key by simply clicking a few times on the website. They can select the desired rate limit per day.

![Image](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/6561960971d956d6cea266da_po_2.png)

### Frontend Development

**Implement the Chat interface**

In the process of creating the UI, we chose libraries and frameworks as follows:

1. **Next.js:** Start a new project using the command: `npx create-next-app my-project`.&#x20;

```bash
npx create-next-app my-project
cd my-project
```

2. **React Hook Form:** Install with: `yarn add react-hook-form @hookform/resolvers zod`. We use Zod for input form validation.

```bash
yarn add react-hook-form @hookform/resolvers zod
```

3. **TanStack Query:** Install TanStack Query (React Query) in the project: `yarn add react-query`.

```bash
yarn add react-query
```

**Connect the frontend with the backend**

Now, we will implement the Chat Interface with a Streamed Text approach to receive real-time text data from the API and display it.

```javascript
codeuseEffect(() => {
  const message = { query: watch('query') };
  const getData = async () => {
    try {
      setValue('query', '');
      const response = await fetch(
        `${API_BOT}/query?uuid=${localStorage?.session}&message=${message.query}`,
        {
          method: 'GET',
          headers: {
            Accept: 'text/event-stream',
            'x-api-key': localStorage?.apiKey, // API key
          },
        }
      );
      const reader = response.body!.getReader();
      let result = '';
      while (true) {
        const { done, value } = await reader?.read();
        if (done) {
          setStreamText('');
          setAnswer((prevState) => [
            ...prevState,
            {
              id: (prevState.length + 1).toString(),
              role: 'ai',
              message: result,
            },
          ]);
          break;
        }
        result += new TextDecoder().decode(value);
        setStreamText(
          (prevData) => prevData + new TextDecoder().decode(value)
        );
      }
    } catch (error: any) {
      console.error(error);
      setError('bot', {
        message: error?.response?.data?.message ?? 'Something went wrong',
      });
    }
  };
  if (isSubmitSuccessful) {
    getData();
  }
}, [submitCount, isSubmitSuccessful, setValue, watch, setError]);
```

This code snippet fetches data from the API using the Streamed Text approach, allowing real-time updates of chat messages.

**Implementing API key generation**

After setting up the UI interface, we connect it to the backend for API key generation.

```json
// API Plan
// GET {endpoint}/plan
[
  {
    "name": "20RequestPerDay"
  },
  {
    "name": "30RequestPerDay"
  },
  {
    "name": "10RequestPerDay"
  }
]

// API Create Key
// POST {endpoint}/create_key
// Body
{
  "plan_name": "30RequestPerDay",
  "user": "example@email.com"
}

// Response 200 (OK)
{
  "value": "7wuOmY7Osz721X1bQUkAP2aGUF5oOVw28EJx7MVS"
}
```

After obtaining the API key, we use it in other projects by adding the header:

```json
{
  "x-api-key": "7wuOmY7Osz721X1bQUkAP2aGUF5oOVw28EJx7MVS"
}
```

**Implementing rate limit notification**

We handle input validation errors and errors from API usage in the Stream. When the API usage exceeds the limit, the API will return a status code of 429. In this case, we set an error message to notify the user.

```javascript
if (response.status === 429) {
  setError('bot', {
    message: 'Limit Exceeded',
  });
  return;
}
```

This ensures that the UI displays an error message when the user exceeds the API usage limit.

### Backend Development

#### Configuring Usage Plans

Begin by creating the usage plans that you want. Create three plans: 10RequestPerDay, 20RequestPerDay, and 30RequestPerDay. Map these plans to the API Gateway you created, indicating which plans should be used by which API Gateway.

#### Configuring API Gateway to Use API Key

Configure each method in API Gateway to require an API key. Edit the method request in the desired resource and check the "API Key required" option.

#### Setting up FastAPI Backend

Install the necessary Python libraries for FastAPI:

```bash
pip install fastapi
pip install "uvicorn[standard]"
```

Create a file named `app.py` and initialize the FastAPI application:

```python
pythonCopy codefrom fastapi import FastAPI
from fastapi.encoders import jsonable_encoder
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware

app = FastAPI()

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

@app.get("/helloworld")
async def helloworld():
    return {"message": "Hello World"}
```

This code initializes a FastAPI project, enables CORS for frontend communication, and sets up a simple endpoint. Run the server using:

```bash
uvicorn app:app
```

#### Connecting to AWS with boto3

Use the boto3 library to interact with AWS services. Make sure that the EC2 instance has sufficient permissions for the intended actions.

```bash
pip install boto3
```

Create a boto3 client for API Gateway:

```python
import boto3

client = boto3.client(
    'apigateway',
    region_name='ap-southeast-1'
)
```

#### API Design and Deployment

**Design**

Create two API endpoints:

1. **To view available usage plans:**

```python
def get_plan():
    res = client.get_usage_plans()
    plan_name_list = []
    for i in res['items']:
        plan_name_list.append({
            "name": i['name']
        })
    return JSONResponse(content=plan_name_list, status_code=200)
```

2. **To create an API key and associate it with a usage plan:**

```python
from fastapi import HTTPException

@app.post("/create_key")
async def create_key(create_request: create_request):
    res = module.check_user_exist_in_plan(create_request.user, create_request.plan_name)
    if res:
        raise HTTPException(status_code=409, detail="Key already exists")
    else:
        response = client.create_api_key(
            name=create_request.user,
            description='for use api',
            enabled=True,
            generateDistinctId=True
        )
        api_key = response['value']
        api_id = response['id']
        try:
            module.add_key_to_plan(api_id, create_request.plan_name)
            result = {
                "value": api_key
            }
            return JSONResponse(content=result, status_code=200)
        except Exception as e:
            client.delete_api_key(
                apiKey=api_id
            )
            result = {
                "message": "create failed"
            }
            return JSONResponse(content=result, status_code=500)
```

This endpoint creates an API key, associates it with the specified usage plan, and returns the generated API key.

#### Deploy

Use the `uvicorn` command to run the FastAPI server:

```
uvicorn app:app --host 0.0.0.0
```

This command runs the server on port 8000.

#### Testing API Calls with API Key

To test API calls with an API key through AWS API Gateway, include the key in the header as `x-api-key`:

```json
{
  "x-api-key": "7wuOmY7Osz721X1bQUkAP2aGUFxxxxxxxx"
}
```

This key will be read by API Gateway and sent to the backend, which you configured earlier.


# OpenAI with Guardrail

#### **Use Case Overview**

We have created a chatbot, dividing it into two sides: the user side and the admin side. This is a simulation where the user interacts with the chatbot by asking questions, and the admin is the owner or the one who defines the criteria for the chatbot's questions.

Users can ask and answer questions just like the previous chatbot we created. However, there is a difference; they can only ask questions within the specified criteria. For example, this chatbot is set to only respond to questions related to animals. If the user asks a question unrelated to animals, the chatbot will reply that it cannot answer that question.

On the admin side, there is the privilege to set the criteria for the number of questions, and it is possible to add and clear criteria in the admin section of the chatbot, which displays all the criteria currently in use.

### User Flow

#### **Admin**

![Admin Interface for Rule Management](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/656a2e7928457dee4ce1d53c_PO_3.png)

#### **User**

![User Interface for Question Submission](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/656a2e80bc0ada798aebfc8a_PO_4.png)

### Frontend Development

#### **Admin Interface for Rule Management**

In designing the UI for the admin's rule management, the admin can create and delete rules and view the rules set for development. For the admin interface, React is used to manage various states. For form management, the library `react-hook-form`, `@hookform/resolvers/zod`, and `zod` are used for input validation. The connection to the API is handled using `@tanstack/react-query`, which provides various features for API interaction, such as using `useQueries` for fetching multiple values simultaneously.

```typescript
const [{ data: session }, { data: rule, refetch: refetchRule }] = useQueries({
    queries: [
      {
        queryKey: ['session', endPoint],
        queryFn: async () => {
          const { data } = await axios.get(`/session`)
          return data
        },
        enabled: typeof window !== 'undefined' && !localStorage?.session,
      },
      {
        queryKey: ['rule', endPoint],
        queryFn: async () => {
          const { data } = await axios.get(`/rule`)
          return data
        },
      },
    ],
  })
```

For adding and clearing rules, `useMutation` is used. The advantage of this library is that it handles various states automatically. For example, after successfully adding or clearing a rule, it triggers the `refetchRule` function to update the table.

```typescript
const { mutateAsync: createRule } = useMutation({
    mutationFn: (rule: string) => {
      return axios.post('/create_rule', { rule })
    },
    onSuccess: () => {
      refetchRule()
    },
  })

  const { mutateAsync: clearRule } = useMutation({
    mutationFn: () => {
      return axios.get('/clear_collection')
    },
    onSuccess: () => {
      refetchRule()
    },
  })
```

#### **User Interface for Question Submission**

In the implementation of the chatbot interface, there is an input section for typing messages to the AI. When a user types something unrelated to the specified rules, the API responds that it is a "bad prompt" and prompts the user to ask another question.

Now, let's see what happens when we ask a question related to cats. The AI responds by asking about the cat we inquired about. The interaction with the chatbot is done in a streaming format.

```typescript
useEffect(() => {
    const message = { query: watch('query') }
    const getData = async () => {
      try {
        setValue('query', '')
        const response = await fetch(
          `${endPoint ?? API_URL}/query?uuid=${localStorage?.session}&message=${
            message.query
          }`,
          {
            method: 'GET',
            headers: {
              Accept: 'text/event-stream',
              'x-api-key': localStorage?.apiKey,
            },
          }
        )

        if (response.status === 200) {
          const reader = response.body!.getReader()
          let result = ''
          while (true) {
            const { done, value } = await reader?.read()
            if (done) {
              setStreamText('')
              setAnswer((prevState) => [
                ...prevState,
                {
                  id: (prevState.length + 1).toString(),
                  role: 'ai',
                  message: result,
                },
              ])
              break
            }
            result += new TextDecoder().decode(value)
            setStreamText(result)
          }
        } else {
          if (response.status === 404) {
            localStorage.removeItem('session')
            refreshSession()
            setError('bot', {
              message: 'Something went wrong, please try again',
            })
            return
          }

          setError('bot', {
            message: 'Something went wrong',
          })
        }
      } catch (error: any) {
        console.error(error)
        setError('bot', {
          message: error?.response?.data?.message ?? 'Something went wrong',
        })
      }
    }
    if (isSubmitSuccessful && localStorage) {
      getData()
    }
  }, [
    submitCount,
    isSubmitSuccessful,
    setValue,
    watch,
    setError,
    endPoint,
    refreshSession,
  ])
```

#### **Guardrail Interaction**

The Guardrail feature limits the chatbot's ability to respond to prompts based on the rules set by the admin. The Relevant Answer Generation (RAG) checks how closely the prompt matches the rules and sends it to OpenAI. If the similarity is less than 99%, the API responds with a "bad prompt" message, indicating that the question is outside the defined rules.

#### **Error Handling and User Feedback**

Error handling is implemented for server connection issues or user input errors. For example, if there's a 404 status code (not found), the session is cleared and a new session is fetched. Validation errors for user input are handled using `react-hook-form`, `@hookform/resolvers/zod`, and `zod`.

```typescript
export const askScheme = z.object({
  apiKey: z.string().optional(),
  query: z.string().trim().min(1, { message: 'Please enter your message' }),
  bot: z.string({}).optional(),
})

export const ruleScheme = z.object({
  rule: z.string().trim().min(1, { message: 'Please enter your rule' }),
})

export const endpointScheme = z.object({
  endpoint: z.string().url().optional(),
})

const methodRule = useForm<IRuleForm>({
    resolver: zodResolver(ruleScheme),
    mode: 'onChange',
    shouldFocusError: true,
  })

  const methodsEndpoint = useForm<IEndPointForm>({
    resolver: zodResolver(endpointScheme),
    mode: 'onChange',
    shouldFocusError: true,
  })

  const methods = useForm<IOpenAIForm>({
    resolver: zodResolver(askScheme),
    mode: 'onChange',
    shouldFocusError: true,
    defaultValues: {
      apiKey: '',
      query: '',
    },
  })
```

[Link to Frontend Code](https://github.com/vultureprime/ai-web-interface/tree/main/next-open-ai-guardrail)

### Backend Development

#### **Setup FastAPI**

Install the necessary Python 3 libraries for FastAPI:

```python
pip install fastapi
pip install "uvicorn[standard]"
```

Create a file named `guardrail.py` and initialize FastAPI:

```python
from fastapi import FastAPI
from fastapi.encoders import jsonable_encoder
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware

app = FastAPI()

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

@app.get("/helloworld")
async def helloworld():
    return {"message": "Hello World"}
```

Run the FastAPI server using the command:

```python
uvicorn guardrail:app
```

This command instructs FastAPI, declared in the `app.py` file, to run, with the server defaulting to port 8000.

The API is divided into two parts: one for users sending prompts to the backend for interacting with AI, and another for administrators controlling whether prompts can be sent to AI.

Install other necessary dependencies:

```python
pip install openai
pip install uuid
pip install pydantic
pip install langchain
```

The admin API includes functions for creating, viewing, and deleting rules for filtering prompts, essentially adding data to VectorDB. The VectorDB initialization will be discussed further.

```python
def add_rule(text):
    Chroma.from_texts(collection_name=collection_name,texts=[text], embedding=embedding ,persist_directory=persist_dir)

@app.post('/create_rule')
def create_rule(rule:Rule):
    add_rule(rule.rule)
    return JSONResponse(content={
        "message":"success",
        "rule":rule.rule
        })
```

View all created rules:

```python
def get_collection():
    vectorstore = init_db()
    result = vectorstore.get()
    return result['documents']

@app.get('/rule')
def get_rule():
    res = get_collection()
    return JSONResponse(content={"result": res})
```

Clear all rules:

```python
def clear_db_collection():
    vectorstore = init_db()
    res = vectorstore.delete_collection()

@app.get('/clear_collection')
def clear_collection():
    clear_db_collection()
    return JSONResponse(content={"message": "remove complete"})
```

For user interaction, an API is provided to send prompts and check whether they match any rules:

```python
@app.get("/session")
def session():
    client_uuid = uuid.uuid4()
    create_session(str(client_uuid))
    result = {
        "uuid" : str(client_uuid)
    }
    return JSONResponse(content=result)

@app.get("/query")
async def main(uuid:str,message:str):
    get_session(uuid)
    collection = get_collection()
    if  len(collection) == 0:
        return StreamingResponse(
                    generate_response('No rules configure please ask admin'), 
                    media_type="text/event-stream"
                )
    else:
        score = compare_similarity(message)
        print(score)
        if score > 0.7:
            return StreamingResponse(
                        stream_chat(
                            uuid = uuid,
                            prompt= message
                        ), 
                        media_type="text/event-stream"
                    )
        else:
            return StreamingResponse(
                        generate_response('bad prompt please ask another one'), 
                        media_type="text/event-stream"
                    )
```

#### **Integrating OpenAI API**

Connect to the OpenAI API using the API key generated from the OpenAI console. Two methods are provided: one using an Embedding model and the other using the Chatbot API.

1. Embedding model:

```python
from langchain.embeddings import OpenAIEmbeddings
api_key = 'sk-XcTnjgYVsJQMNxxxxxxxxxxxxxx'
embedding = OpenAIEmbeddings(openai_api_key=api_key)
```

2. Chatbot API:

```python
import openai
openai.api_key = api_key

def stream_chat(uuid: str, prompt: str):
    result = ""
    messages = add_message(uuid, 'user', prompt)
    for chunk in openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=messages,
        stream=True,
    ):
        content = chunk["choices"][0].get("delta", {}).get("content")
        if content is not None:
            result = result + content
            yield content
    add_message(uuid, 'assistant', result)
```

#### **Implementing VectorDB with Chroma**

Chroma is used for VectorDB, which stores vectorized data to aid in finding similar data. Initialization is done using parameters for collection name, data location, and the embedding model.

```python
from langchain.vectorstores import Chroma

def init_db():
    vectorstore = Chroma(collection_name=collection_name, persist_directory=persist_dir, embedding_function=embedding)
    return vectorstore
```

#### **Building the Guardrail Feature**

The Guardrail feature queries rule data from VectorDB based on received prompts, and if the similarity score is above a threshold, the prompt is forwarded to the chatbot.

```python
def compare_similarity(query):
    vectorstore = init_db()
    result = vectorstore.similarity_search_with_relevance_scores(query, k=5)
    score_list = []
    for i in result:
        score_list.append(i[-1])
    try:
        average = sum(score_list)/len(score_list)
        return average
    except:
        return 0
```

#### **Deploying on AWS EC2**

Deployment is done using the `screen` utility:

1. Create a session with a specific name:

```bash
screen -S name
```

2. Navigate to the API folder:

```bash
cd path/to/api
```

3. Start FastAPI on port 8000:

```bash
uvicorn guardrail:app --host 0.0.0.0 --port 8000
```

4. Detach from the current screen session:

```bash
Ctrl+a d
```

#### **API Gateway and CORS Configuration**

Add CORS configuration for FastAPI:

```python
from fastapi.middleware.cors import CORSMiddleware

app.app = FastAPI()
app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)
```

Connect this server to AWS API Gateway for authentication and usage management in each request.

[Link to Backend Code](https://github.com/vultureprime/ai-web-backend/tree/main/Guardrail-openai)


# Multiple Agents

#### Use Case Overview

We have developed a Chatbot specialized in website development. The AI system is designed with the first AI acting as a task divider for different agents, including Frontend, Backend, and Designer. Each agent responds to questions in its respective domain. Additionally, there is another AI responsible for summarizing responses from all agents.

When a user asks a question, such as "create blog," the response includes information from Frontend, Backend, and Designer, presented in clear sections for easy comprehension.

### User Flow

![User Flow](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/657462924ee49de84a8b755d_po_2.png)

#### AI

In the realm of Large Language Models (LLMs), we continue to use OpenAI as the model for creating the Chatbot. However, we have separated it into multiple agents, each with a distinct role.

**Manager Agent**

The initial AI is the Manager, with two main responsibilities. Firstly, it manages the segregation of prompts from users, assigning suitable tasks to each agent. This enhances our ability to handle user requests effectively. Secondly, it collects answers, summarizes results, and organizes data from Task Processing Agents before presenting the response to the user.

**Task Processing Agent**

In the Task Processing segment, or AI responsible for generating responses, we have divided it into three agents:

* Frontend: Generates responses for questions related to Frontend.
* Backend: Handles responses in the Backend and Technical domain.
* Designer: Generates responses related to UI and UX design.

These agents work based on the tasks assigned by Manager Agents. Once their tasks are complete, they send the responses back to the Manager Agent for further consolidation.

#### API

For the API, we continue to use FastAPI as the framework. The demo API includes the following:

**/query**

We have created an API that combines the usage of all agents in a single call. Frontend can call this API once when a user asks a question. Subsequently, it generates responses through various agents and displays the results in the chat without the need for multiple API calls.

The APIs called within this API include:

* **/breakdown**: Used to break down questions into tasks for forwarding to specific agents.
* **/build**: Used to generate responses.
* **/conclude**: Used to collect responses and summarize them for presentation to the user.

**/queryWithOutChain**

For users who do not want to use multiple agents, we provide an API to connect directly to OpenAI.

### Frontend Development

#### **Implementing Input Box for User Prompts**&#x20;

To implement the input, we will use the `react-hook-form` library along with `@hookform/resolvers/zod` and 'zod'. In the first step, we create a schema comprising a query for input and a bot for displaying error messages related to the bot itself.

```typescript
export const askScheme = z.object({
  query: z.string().trim().min(1, { message: 'Please enter your message' }),
  bot: z.string({}).optional(),
});
```

Once the schema is created, we generate a type interface for this form.

```typescript
export interface IOpenAIForm extends z.infer<typeof askScheme> {}
```

Next, we create methods:

```typescript
const methods = useForm<IOpenAIForm>({
    resolver: zodResolver(askScheme),
    shouldFocusError: true,
    defaultValues: {
      query: '',
    },
  });

const { handleSubmit, setError, setValue } = methods;

const onSubmit = async (data: IOpenAIForm) => {
  // ...TO DO Something
};

return (
  <FormProvider methods={methods} onSubmit={handleSubmit(onSubmit)}>
    {/* ... */}
  </FormProvider>
);
```

The schema validates user input, and if the user submits without typing anything, an error message will be displayed in the user interface: 'Please enter your message.' The input component uses the 'react-hook-form' library along with `useFormContext` and `Controller` to manage input.

```typescript
import { InputHTMLAttributes } from 'react';
import { useFormContext, Controller } from 'react-hook-form';
import { twMerge } from 'tailwind-merge';

interface IProps extends InputHTMLAttributes<HTMLInputElement> {
  name: string;
  helperText?: string;
  label?: string;
}

const RHFTextField = ({
  name,
  helperText,
  label,
  className,
  ...other
}: IProps) => {
  const { control } = useFormContext();

  return (
    <Controller
      name={name}
      control={control}
      render={({ field, fieldState: { error } }) => (
        <div className='w-full flex flex-col gap-y-2 '>
          {label && <label className='text-sm font-semibold'>{label}</label>}
          <input
            {...field}
            {...other}
            className={twMerge(
              'outline-none w-full  border border-gray-200 rounded-lg px-2 py-1',
              className
            )}
          />
          {(!!error || helperText) && (
            <div className={twMerge(error?.message && 'text-rose-500 text-sm')}>
              {error?.message || helperText}
            </div>
          )}
        </div>
      )}
    />
  );
};

export default RHFTextField;
```

When using the input, set the `name` attribute to 'query' to match the variable declared in the schema.

```typescript
const methods = useForm<IOpenAIForm>({
    resolver: zodResolver(askScheme),
    shouldFocusError: true,
    defaultValues: {
      query: '',
    },
  });

const { handleSubmit, setError, setValue } = methods;

const onSubmit = async (data: IOpenAIForm) => {
  console.log(data);
};

return (
  <FormProvider methods={methods} onSubmit={handleSubmit(onSubmit)}>
    {/* ... */}
    <RHFTextField
      type='text'
      placeholder='What do you need? ...'
      className='outline-none w-full border-none'
      name='query'
    />
    <button
      type='submit'
      disabled={isLoading}
      className='text-gray-400 disabled:text-gray-200'
    >
      Submit
    </button>
  </FormProvider>
);
```

When the submit button is pressed, if the input is validated correctly, we see the logged data from the `onSubmit` function. This data can then be sent to the backend for further processing.

#### **Handle API Integration for Chained-LLM Agents & Display Output from Each API**&#x20;

To connect with the backend, we manage various states such as loading and error. We set the default base URL for the API.

```typescript
axios.defaults.baseURL = process.env.NEXT_PUBLIC_API;
```

After setting the base URL, we connect to the API.

```typescript
const onSubmit = async (data: IOpenAIForm) => {
  try {
    const { data: result } = await axios.post(
      `/query?question=${data.query}`,
      undefined
    );
    console.log(result); // Value obtained from the API
    // TO DO Something
  } catch (error) {
    const err = error as AxiosError<{ detail: string }>;
    setError('bot', {
      message: err?.response?.data?.detail ?? 'Something went wrong',
    });
  }
};
```

We use `useFormContext` from 'react-hook-form' to manage the loading and errors state.

```typescript
const {
  formState: { isSubmitting, errors },
} = useFormContext();
```

If an API error occurs, we display an error message from the bot and show the loading state. We use the state from 'useFormContext' to display values such as `isSubmitting` and `errors`.

```typescript
export default function ChatWidget({ answer }: { answer: ChatProps[] }) {
  const chatWindowRef = useRef<HTMLDivElement>(null);
  const {
    formState: { isSubmitting, errors },
  } = useFormContext();

  return (
    <div className='h-full flex flex-col w-full'>
      <Header />
      <ChatWindow
        messages={answer}
        isLoading={isSubmitting}
        error={errors?.bot?.message as string}
        chatWindowRef={chatWindowRef}
      />
      <ChatInput isLoading={isSubmitting} />
    </div>
  );
}
```

In the `ChatWindow` component, we manage various states to display in the UI, including the results from the API.

```typescript
import { useEffect } from 'react'
import Image from 'next/image'
import { CopyClipboard } from './CopyClipboard'

interface Message {
  role: 'user' | 'ai'
  message: string
  id: string
  raw: string
}

interface ChatWindowProps {
  messages: Message[]
  isLoading?: boolean
  error?: string
  chatWindowRef: any | null
}

export const ChatWindow: React.FC<ChatWindowProps> = ({
  messages,
  isLoading,
  error,
  chatWindowRef,
}) => {
  useEffect(() => {
    if (
      chatWindowRef !== null &&
      chatWindowRef?.current &&
      messages.length > 0
    ) {
      chatWindowRef.current.scrollTop = chatWindowRef.current.scrollHeight
    }
  }, [messages.length, chatWindowRef])

  return (
    <divref={chatWindowRef}
      className='flex-1 overflow-y-auto p-4 space-y-8'
      id='chatWindow'
    >
      {messages.map((item, index) => (
        <div key={item.id} className='w-full'>
          {item.role === 'user' ? (
            <div className='flex gap-x-8 '>
              <div className='min-w-[48px] min-h-[48px]'>
                <Imagesrc='/img/chicken.png'
                  width={48}
                  height={48}
                  alt='user'
                />
              </div>
              <div>
                <p className='font-bold'>User</p>
                <p>{item.message}</p>
              </div>
            </div>
          ) : (
            <div className='flex gap-x-8 w-full'>
              <div className='min-w-[48px] min-h-[48px]'>
                <Imagesrc='/img/robot.png'
                  width={48}
                  height={48}
                  alt='robot'
                />
              </div>
              <div className='w-full'>
                <div className='flex justify-between mb-1 w-full '>
                  <p className='font-bold'>Ai</p>
                  <div />
                  <CopyClipboard content={item.raw} />
                </div>

                <divclassName='prose whitespace-pre-line'
                  dangerouslySetInnerHTML={{ __html: item.message }}
                />
              </div>
            </div>
          )}
        </div>
      ))}
      {isLoading && (
        <div className='flex gap-x-8 w-full mx-auto'>
          <div className='min-w-[48px] min-h-[48px]'>
            <Image src='/img/robot.png' width={48} height={48} alt='robot' />
          </div>
          <div>
            <p className='font-bold'>Ai</p>

            <div className='mt-4 flex space-x-2 items-center '>
              <p>Hang on a second </p>
              <span className='sr-only'>Loading...</span>
              <div className='h-2 w-2 bg-blue-600 rounded-full animate-bounce [animation-delay:-0.3s]'></div>
              <div className='h-2 w-2 bg-blue-600 rounded-full animate-bounce [animation-delay:-0.15s]'></div>
              <div className='h-2 w-2 bg-blue-600 rounded-full animate-bounce'></div>
            </div>
          </div>
        </div>
      )}
      {error && (
        <div className='flex gap-x-8 w-full mx-auto'>
          <div className='min-w-[48px] min-h-[48px]'>
            <Image src='/img/error.png' width={48} height={48} alt='error' />
          </div>
          <div>
            <p className='font-bold'>Ai</p>
            <p className='text-rose-500'>{error}</p>
          </div>
        </div>
      )}
    </div>
  )
}
```

#### **Implement Process Continuation Functionality**&#x20;

After successfully connecting with the backend, we need to store the result message obtained from the API. We create a state to store the user and bot answers.

```typescript
export default function ChatBotDemo() {
  const [answer, setAnswer] = useState<ChatProps[]>([])

  const methods = useForm<IOpenAIForm>({
    resolver: zodResolver(askScheme),
    shouldFocusError: true,
    defaultValues: {
      query: '',
    },
  })

  const { handleSubmit, setError, setValue } = methods

  const onSubmit = async (data: IOpenAIForm) => {
    try {
      const id = answer.length
      setAnswer((prevState) => [
        ...prevState,
        {
          id: id.toString(),
          role: 'user',
          message: data.query,
          raw: '',
        },
      ])
      setValue('query', '')
      const { data: result } = await axios.post(
        `/query?question=${data.query}`,
        undefined
      )
      setAnswer((prevState) => [
        ...prevState,
        {
          id: (prevState.length + 1).toString(),
          role: 'ai',
          message: result.raw,
          raw: result.json.customer_need,
        },
      ])
    } catch (error) {
      const err = error as AxiosError<{ detail: string }>
      setError('bot', {
        message: err?.response?.data?.detail ?? 'Something went wrong',
      })
    }
  }

  return (
    <FormProvider methods={methods} onSubmit={handleSubmit(onSubmit)}>
      <div className='flex justify-center flex-col items-center bg-white mx-auto max-w-7xl h-screen '>
        <ChatWidget answer={answer} />
      </div>
    </FormProvider>
  )
}
```

Now, when we ask the AI, for example, to create a blog, the LLMs agent will respond with tasks for each role, such as Design, Frontend, and Backend.

[Link to Frontend Code](https://github.com/vultureprime/ai-web-interface/tree/main/next-llms-agent)

### Backend Development

#### Setting Up a FastAPI Project

Similar to all the examples we've gone through, we choose to use FastAPI as the framework to create and use our API.

Start by installing the necessary Python libraries for FastAPI:

```bash
pip install fastapi
pip install "uvicorn[standard]"
```

Create a file named `app.py` and initialize FastAPI:

```python
from fastapi import FastAPI
from fastapi.encoders import jsonable_encoder
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware

app = FastAPI()

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

@app.get("/helloworld")
async def helloworld():
    return {"message": "Hello World"}
```

In this example code, we initialize a FastAPI project and enable CORS for connecting with the frontend. To run the server, use the following command:

```bash
uvicorn app:app
```

This command instructs FastAPI, declared in the `app.py` file, to work. The server runs on the default port 8000.

Now, let's create another file named `LocalTemplate.py` to store initial templates for asking questions to the Chatbot. These templates include:

* `Manager-template`: Divides tasks for incoming questions to be used by each agent.
* `Agent-template`: Divides into three agents: Frontend, Backend, and Designer. Each agent answers questions related to their role.
* `Conclusion-template`: Summarizes all received answers for a concise overview.

In the API design section, we'll divide the API into five routes for different functionalities, which will be explained in the next section.

#### Developing the Manager Agent API

The first API we create is `POST: /breakdown`, which handles the breakdown of a customer's need into tasks for each agent:

```python
@app.post('/breakdown')
def breakdown_question(customer_need : str):
    model_256 = ChatOpenAI(model_name="gpt-4-1106-preview", temperature=0.3, max_tokens=256,openai_api_key = OPENAI_API_KEY)
    breakdown_chain = ChatPromptTemplate.from_template(LocalTemplate.get_manager()) | model_256
    result = breakdown_chain.invoke({"question": customer_need})
    arr = result.content.split('<question>')[1:]
    task_list = []
    for i in arr : 
        full_task = remove_tag(i,['<question>','<role>','</question>','</role>','</sub-question>']).strip()
        full_task_list = full_task.split('<sub-question>')
        role = full_task_list[0]
        task = full_task_list[1]
        task_list.append({'role':role,'task':task})

    json_compatible_item_data = jsonable_encoder(task_list)
    return JSONResponse(content=json_compatible_item_data)
```

This API takes a customer's need as input and processes it to create tasks for each agent. The resulting tasks are then returned as a JSON response.

#### Building Task Processing APIs

Next, we create an API `POST: /build` that processes the tasks for each agent:

```python
def build_task(task_list : List[task_list]):
    model_256 = ChatOpenAI(model_name="gpt-4-1106-preview", temperature=0.3, max_tokens=256,openai_api_key = OPENAI_API_KEY)
    frontend_chain = ChatPromptTemplate.from_template(LocalTemplate.get_frontend()) | model_256 
    frontend_result = frontend_chain.invoke({"task": task_list[0].task})
    
    backend_chain = ChatPromptTemplate.from_template(LocalTemplate.get_backend()) | model_256 
    backend_result = backend_chain.invoke({"task": task_list[1].task})

    designer_chain = ChatPromptTemplate.from_template(LocalTemplate.get_designer()) | model_256 
    designer_result = designer_chain.invoke({"task": task_list[2].task})
    
    frontend_text = frontend_result.content.split('<step>')[1:]
    frontend_json = text_to_json(frontend_text,'<description>',['<task>','</task>','<step>','</step>','</description>'])
    backend_text = backend_result.content.split('<step>')[1:]
    backend_json = text_to_json(backend_text,'<description>',['<task>','</task>','<step>','</step>','</description>'])
    designer_text = designer_result.content.split('<step>')[1:]
    designer_json = text_to_json(designer_text,'<description>',['<task>','</task>','<step>','</step>','</description>'])

    result = {
        'raw' : {
            "frontend_task": frontend_result.content, 
            "backend_task": frontend_result.content, 
            "designer_task": frontend_result.content
        },
        'json' : {
            "frontend_task": frontend_json, 
            "backend_task": backend_json, 
            "designer_task": designer_json
        }
    }


    json_compatible_item_data = jsonable_encoder(result)
    return JSONResponse(content=json_compatible_item_data)
```

This API takes a list of tasks and processes them for each agent (Frontend, Backend, Designer), returning the raw and JSON format of the generated tasks.

#### Creating a Manager Summary API

Now, we create an API `POST: /conclude` for summarizing all the received information:

```python
@app.post('/conclude')
def build_conclusion(customer_need : str, frontend_task : str, backend_task : str, designer_task : str):
    model_512 = ChatOpenAI(model_name="gpt-4-1106-preview", temperature=0.3, max_tokens=512,openai_api_key = OPENAI_API_KEY)
    customer_chain = ChatPromptTemplate.from_template(LocalTemplate.get_conclusion()) | model_512
    customer_result = customer_chain.invoke({
        "customer_need" : customer_need,
        "frontend_task": frontend_task, 
        "backend_task": backend_task, 
        "designer_task": designer_task
    })

    customer_json = remove_tag(customer_result.content,['<conclude>','</conclude>','<text>','</text>'])
    result =  {
        'raw' : customer_result.content,
        'json' : {
            'customer_need' : customer_json
        }
    }
    json_compatible_item_data = jsonable_encoder(result)
    return JSONResponse(content=json_compatible_item_data)
```

This API takes the initial customer need and the tasks generated by each agent and produces a summarized conclusion in both raw and JSON formats.

#### Implementing a Chained-LLM Wrapper API

For a more streamlined process, we create an API `POST: /query` that combines all the steps:

```python
@app.post('/query')
def query_with_chain(question : str):
    customer = question
    model_256 = ChatOpenAI(model_name="gpt-4-1106-preview", temperature=0.3, max_tokens=256,openai_api_key = OPENAI_API_KEY)
    model_512 = ChatOpenAI(model_name="gpt-4-1106-preview", temperature=0.3, max_tokens=512,openai_api_key = OPENAI_API_KEY)
    PO_Final_Chain = ChatPromptTemplate.from_template(LocalTemplate.get_manager()) | model_256
    result = PO_Final_Chain.invoke({"question": customer})
    # print(result.content)
    arr = result.content.split('<question>')[1:]
    task_list = []
    for i in arr : 
        full_task = i.replace('\n','').replace('</question>','').replace('<role>','').replace('</role>','').replace('</sub-question>','').strip()
        role = full_task.split('<sub-question>')[0].strip()
        task = full_task.split('<sub-question>')[1].strip()
        task_list.append({'role':role,'task':task})

    frontend_chain = ChatPromptTemplate.from_template(LocalTemplate.get_frontend()) | model_256 
    frontend_result = frontend_chain.invoke({"task": task_list[0]['task']})

    backend_chain = ChatPromptTemplate.from_template(LocalTemplate.get_backend()) | model_256 
    backend_result = backend_chain.invoke({"task": task_list[1]['task']})

    designer_chain = ChatPromptTemplate.from_template(LocalTemplate.get_designer()) | model_256 
    designer_result = designer_chain.invoke({"task": task_list[2]['task']})

    customer_chain = ChatPromptTemplate.from_template(LocalTemplate.get_conclusion()) | model_512
    customer_result = customer_chain.invoke({
        "customer_need" : customer,
        "frontend_task": frontend_result.content, 
        "backend_task": backend_result.content, 
        "designer_task": designer_result.content
    })

    customer_json = remove_tag(customer_result.content,['<conclude>','</conclude>','<text>','</text>'])
    result =  {
        'raw' : customer_result.content,
        'json' : {
            'customer_need' : customer_json
        }
    }

    json_compatible_item_data = jsonable_encoder(result)
    return JSONResponse(content=json_compatible_item_data)
```

This API takes a question, processes it through the entire chain of tasks and agents, and returns the raw and JSON format of the summarized conclusion.

#### Direct OpenAI Prompt API Integration

To have a more direct interaction with the Chatbot, we create an API `POST: /queryWithoutChain`:

```python
def query_without_chain(question : str):
    model_512 = ChatOpenAI(model_name="gpt-4-1106-preview", temperature=0.3, max_tokens=512,openai_api_key = OPENAI_API_KEY)
    chain = ChatPromptTemplate.from_template('{question}') | model_512
    customer_result = chain.invoke({"question": question})
    
    result =  {
        'raw' : customer_result.content,
        'json' : {
            'customer_need' : customer_result.content
        }
    }

    json_compatible_item_data = jsonable_encoder(result)

    return JSONResponse(content=json_compatible_item_data)
```

This API takes a question, sends it directly to the Chatbot without going through the task and agent chain, and returns the raw and JSON format of the Chatbot's response.

#### Deploying and Monitoring on EC2

For deploying FastAPI on an EC2 instance:

1. Create a session with a name of your choice:

   ```bash
   screen -S name
   ```
2. Navigate to the API folder:

   ```bash
   cd path/to/api
   ```
3. Start FastAPI on port 8000:

   ```bash
   uvicorn app:app --host 0.0.0.0
   ```
4. Detach from the screen session:

   ```bash
   Ctrl+a d
   ```

Now, the server runs in the background even if you exit.

#### Setting Up API Gateway and Implementing CORS

To integrate the API with AWS API Gateway:

Add CORS configuration to FastAPI:

```python
from fastapi.middleware.cors import CORSMiddleware

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)
```

This configuration allows CORS for the FastAPI server. Connect this server to AWS API Gateway for enhanced management, authentication, and usage limitation.

[Link to Backend Code](https://github.com/vultureprime/ai-web-backend/tree/main/Agents-openai-langchain)


# Q\&A Chatbots (RAG + Agents)

We have developed a Chatbot consisting of two AI agents and an RAG. When a user asks a question to the Chatbot, AI Agent 1 decides whether to use RAG to search for specific information or use LLMs to answer the question. The Chatbot also includes memory that can remember previous responses, making the interaction smoother.

For example, in the RAG demo, we used metadata containing annual performance data from 2019 to 2022 from Uber. If the user asks a question related to the RAG data, AI Agent 1 decides whether to use RAG or answer the question using OpenAI Agent, depending on the relevance.

In this Chatbot implementation, several engines work together, each serving a different purpose. Let's break down each part:

#### Chat Agent

The Chat Agent uses OpenAI as LLMs and consists of Chat Engine, Sub Question Engine, and RAG Engine. When a user asks a question, the Chat Agent decides whether to answer using RAG or Chat Engine based on the relevance to the RAG metadata description.

#### Sub Question Engine

If RAG is chosen for answering a question, and the prompt requires data from multiple RAG engines, the question is sent to the Sub Question Engine first. It helps break down the question before sending it to the RAG Engine, which is essential since RAG is divided into four sub-engines, each responsible for answering questions related to specific aspects.

#### RAG Engine

The RAG Engine uses performance data from Uber for the years 2019 to 2022. The Chat Agent decides whether to send the question directly to RAG or pass it through the Sub Question Engine based on the relevance to the RAG metadata description.

#### Chat Engine

The Chat Engine, or GPT Engine, answers general questions. When the Chat Agent concludes that the prompt is not related to RAG, it sends the question to the Chat Engine for a general response.

#### Other

* This Chatbot stores conversation data in memory, allowing for smoother interaction by maintaining context.
* To reset the chat or clear the memory, the /resetChat command can be used.
* The /chat endpoint is used for normal queries, while /chatWithoutRAG can be used for queries without involving RAG.

### User Flow

![User Flow](https://assets-global.website-files.com/63cb6b155c56b2dcd14e411d/657d503b89d3663accd7e3a3_PO_2.png)

### Frontend Development

#### **User Interface Components**

The UI components will mainly consist of:

* Chat header: Contains settings for the chat and a button to reset the chat.
* Chat input: Input for typing and sending messages.
* Chat widget: Displays the conversation.

#### **Step 1: Define Schema and Form**

In this step, we will use the library react-hook-form, @hookform/resolvers/zod, and 'zod'.

In the first step, we will create a schema.

The schema consists of a query for input and a bot to display error messages for the bot itself.

```tsx
export const askScheme = z.object({
  query: z.string().trim().min(1, { message: 'Please enter your message' }),
  bot: z.string({}).optional(),
})
```

Once we create the schema, we will create a type interface for this form.

```tsx
export interface IOpenAIForm extends z.infer<typeof askScheme> {}
```

After that, we will create methods.

```tsx
const methods = useForm<IOpenAIForm>({
  resolver: zodResolver(askScheme),
  shouldFocusError: true,
  defaultValues: {
    query: '',
  },
})

const { handleSubmit, setError, setValue } = methods

const onSubmit = async (data: IOpenAIForm) => {
  ...TO DO Something
}

return (
  <FormProvider methods={methods} onSubmit={handleSubmit(onSubmit)}>
    ...
  </FormProvider>
)
```

From the schema itself, we will validate the input received from the user. If the user submits without typing, an error message will be displayed in the user interface saying 'Please enter your message.'

In the input component, the 'react-hook-form' library is used, along with useFormContext and Controller to manage the input.

```tsx
import { InputHTMLAttributes } from 'react'
import { useFormContext, Controller } from 'react-hook-form'
import { twMerge } from 'tailwind-merge'

interface IProps extends InputHTMLAttributes<HTMLInputElement> {
  name: string
  helperText?: string
  label?: string
}

const RHFTextField = ({
  name,
  helperText,
  label,
  className,
  ...other
}: IProps) => {
  const { control } = useFormContext()

  return (
    <Controller name={name}
      control={control}
      render={({ field, fieldState: { error } }) => (
        <div className='w-full flex flex-col gap-y-2 '>
          {label && <label className='text-sm font-semibold'>{label}</label>}
          <input
            {...field}
            {...other}
            className={twMerge(
              'outline-none w-full  border border-gray-200 rounded-lg px-2 py-1',
              className
            )}
          />
          {(!!error || helperText) && (
            <div className={twMerge(error?.message && 'text-rose-500 text-sm')}>
              {error?.message || helperText}
            </div>
          )}
        </div>
      )}
    />
  )
}
export default RHFTextField
```

In the part where the input is called, we will set the attribute name to be 'query,' similar to the variable declared in that schema.

```tsx
const methods = useForm<IOpenAIForm>({
  resolver: zodResolver(askScheme),
  shouldFocusError: true,
  defaultValues: {
    query: '',
  },
})

const { handleSubmit, setError, setValue } = methods

const onSubmit = async (data: IOpenAIForm) => {
  console.log(data)
}

return (
  <FormProvider methods={methods} onSubmit={handleSubmit(onSubmit)}>
      ....
     <RHFTextField type='text'
        placeholder='What do you need ? ...'
        className='outline-none w-full border-none'
        name='query'
      />
      <button type='submit'
        disabled={isLoading}
        className='text-gray-400 disabled:text-gray-200'
      >Submit</button>
  </FormProvider>
)

```

Once we submit, if the input is validated correctly, we will see the data we logged from the onSubmit function, which we can then connect to the backend.

#### **Step 2: Connect Backend**

We will set the default base URL.

```tsx
axios.defaults.baseURL = process.env.NEXT_PUBLIC_API
```

After setting it up, we will connect to the API.

```tsx
const onSubmit = async (data: IOpenAIForm) => {
  try {

    const { data: result } = await axios.post(
      `/chat`,
      { query:data.query },
    )
    console.log(result) //Value obtained from API
    //TO DO Something
  } catch (error) {
    const err = error as AxiosError<{ detail: string }>
    setError('bot', {
      message: err?.response?.data?.detail ?? 'Something went wrong',
    })
  }
}
```

When we try to submit the form, we will get the value from the API

We will then connect the obtained data to the UI of the Chatbot using React and State Management.

We will use **`useState`** from React for state management of messages to display in the UI. We will have **`answer`** and **`setAnswer`** to store the questions and answers of the user and bot. The structure of the array will be as follows:

```tsx
[
    {
        "id": "0",
        "role": "user",
        "message": "What were some of the biggest risk factors in 2022 for Uber?",
        "raw": ""
    },
    {
        "id": "2",
        "role": "ai",
        "message": "Some of the biggest risk factors for Uber in 2022 include:\\n\\n1. Reclassification of drivers: There is a risk that drivers may be reclassified as employees or workers instead of independent contractors. This could result in increased costs for Uber, including higher wages, benefits, and potential legal liabilities.\\n\\n2. Intense competition: Uber faces intense competition in the mobility, delivery, and logistics industries. Competitors may offer similar services at lower prices or with better features, which could result in a loss of market share for Uber.\\n\\n3. Need to lower fares or service fees: To remain competitive, Uber may need to lower fares or service fees. This could impact the company's revenue and profitability.\\n\\n4. Significant losses: Uber has incurred significant losses since its inception. The company may continue to experience losses in the future, which could impact its financial stability and ability to attract investors.\\n\\n5. Uncertainty of achieving profitability: There is uncertainty regarding Uber's ability to achieve or maintain profitability. The company expects operating expenses to increase, which could make it challenging to achieve profitability in the near term.\\n\\nThese risk factors highlight the challenges and uncertainties that Uber faces in 2022.",
        "raw": "Some of the biggest risk factors for Uber in 2022 include:\\n\\n1. Reclassification of drivers: There is a risk that drivers may be reclassified as employees or workers instead of independent contractors. This could result in increased costs for Uber, including higher wages, benefits, and potential legal liabilities.\\n\\n2. Intense competition: Uber faces intense competition in the mobility, delivery, and logistics industries. Competitors may offer similar services at lower prices or with better features, which could result in a loss of market share for Uber.\\n\\n3. Need to lower fares or service fees: To remain competitive, Uber may need to lower fares or service fees. This could impact the company's revenue and profitability.\\n\\n4. Significant losses: Uber has incurred significant losses since its inception. The company may continue to experience losses in the future, which could impact its financial stability and ability to attract investors.\\n\\n5. Uncertainty of achieving profitability: There is uncertainty regarding Uber's ability to achieve or maintain profitability. The company expects operating expenses to increase, which could make it challenging to achieve profitability in the near term.\\n\\nThese risk factors highlight the challenges and uncertainties that Uber faces in 2022."
    }
]

```

We will also manage state for shooting the API. When we ask, we can choose whether to use RAG or not. We will have **`hasRag`** and **`setHasRag`** to manage the state, allowing us to use this value to check before sending the API to decide which one to shoot.

```tsx
import ChatWidget, { ChatProps } from '@/app/components/ChatWidget'
import FormProvider from '@/app/components/hook-form/FormProvider'
import { zodResolver } from '@hookform/resolvers/zod'
import axios, { AxiosError } from 'axios'
import { useState } from 'react'
import { useForm } from 'react-hook-form'
import { z } from 'zod'

export const askScheme = z.object({
  query: z.string().trim().min(1, { message: 'Please enter your message' }),
  bot: z.string({}).optional(),
})

export enum ChatType {
  Basic,
  WithoutRag,
}

axios.defaults.baseURL = process.env.NEXT_PUBLIC_API

export interface IOpenAIForm extends z.infer<typeof askScheme> {}

export default function ChatBotDemo() {
  const [answer, setAnswer] = useState<ChatProps[]>([])
  const [hasRag, setHasRag] = useState(true)

  const methods = useForm<IOpenAIForm>({
    resolver: zodResolver(askScheme),
    shouldFocusError: true,
    defaultValues: {
      query: '',
    },
  })

  const { handleSubmit, setError, setValue } = methods

  const onSubmit = async (data: IOpenAIForm) => {
    try {
      const id = answer.length
      setAnswer((prevState) => [
        ...prevState,
        {
          id: id.toString(),
          role: 'user',
          message: data.query,
          raw: '',
        },
      ])
      setValue('query', '')
      const { data: result } = await axios.post(
        `${hasRag ? '/chat' : '/chatWithoutRAG'}`,
        {
          query: data.query,
        }
      )
      setAnswer((prevState) => [
        ...prevState,
        {
          id: (prevState.length + 1).toString(),
          role: 'ai',
          message: result.answer,
          raw: result.answer,
        },
      ])
    } catch (error) {
      const err = error as AxiosError<{ detail: string }>
      setError('bot', {
        message: err?.response?.data?.detail ?? 'Something went wrong',
      })
    }
  }
  const handleChangeRag = () => {
    setHasRag(!hasRag)
  }

  return (
    <FormProvider methods={methods} onSubmit={handleSubmit(onSubmit)}>
      <div className='flex justify-center flex-col items-center bg-white mx-auto max-w-7xl h-screen '>
        <ChatWidget answer={answer} option={{ hasRag, handleChangeRag }} />
      </div>
    </FormProvider>
  )
}
```

#### Step 3: Handle Form Submission

In this step, we handle various states such as loading, submitting, and errors. We use state from 'useFormContext' to display values related to the form, including `isSubmitting` and `errors`.

```tsx
export default function ChatWidget({ answer }: { answer: ChatProps[] }) {
  const chatWindowRef = useRef<HTMLDivElement>(null)
  const {
    formState: { isSubmitting, errors },
  } = useFormContext()

  return (
    <div className='h-full flex flex-col w-full'>
      <Header />
      <ChatWindow messages={answer} isLoading={isSubmitting} error={errors?.bot?.message as string} chatWindowRef={chatWindowRef} />
      <ChatInput isLoading={isSubmitting} />
    </div>
  )
}
```

In the `ChatWindow` component, we manage various states to display in the UI, including loading, submitting, and error states.

```tsx
import { useEffect } from 'react'
import Image from 'next/image'
import { CopyClipboard } from './CopyClipboard'

interface Message {
  role: 'user' | 'ai'
  message: string
  id: string
  raw: string
}

interface ChatWindowProps {
  messages: Message[]
  isLoading?: boolean
  error?: string
  chatWindowRef: any | null
}

export const ChatWindow: React.FC<ChatWindowProps> = ({
  messages,
  isLoading,
  error,
  chatWindowRef,
}) => {
  useEffect(() => {
    if (
      chatWindowRef !== null &&
      chatWindowRef?.current &&
      messages.length > 0
    ) {
      chatWindowRef.current.scrollTop = chatWindowRef.current.scrollHeight
    }
  }, [messages.length, chatWindowRef])

  return (
    <divref={chatWindowRef}
      className='flex-1 overflow-y-auto p-4 space-y-8'
      id='chatWindow'
    >
      {messages.map((item, index) => (
        <div key={item.id} className='w-full'>
          {item.role === 'user' ? (
            <div className='flex gap-x-8 '>
              <div className='min-w-[48px] min-h-[48px]'>
                <Imagesrc='/img/chicken.png'
                  width={48}
                  height={48}
                  alt='user'
                />
              </div>
              <div>
                <p className='font-bold'>User</p>
                <p>{item.message}</p>
              </div>
            </div>
          ) : (
            <div className='flex gap-x-8 w-full'>
              <div className='min-w-[48px] min-h-[48px]'>
                <Imagesrc='/img/robot.png'
                  width={48}
                  height={48}
                  alt='robot'
                />
              </div>
              <div className='w-full'>
                <div className='flex justify-between mb-1 w-full '>
                  <p className='font-bold'>Ai</p>
                  <div />
                  <CopyClipboard content={item.raw} />
                </div>

                <divclassName='prose whitespace-pre-line'
                  dangerouslySetInnerHTML={{ __html: item.message }}
                />
              </div>
            </div>
          )}
        </div>
      ))}
      {isLoading && (
        <div className='flex gap-x-8 w-full mx-auto'>
          <div className='min-w-[48px] min-h-[48px]'>
            <Image src='/img/robot.png' width={48} height={48} alt='robot' />
          </div>
          <div>
            <p className='font-bold'>Ai</p>

            <div className='mt-4 flex space-x-2 items-center '>
              <p>Hang on a second </p>
              <span className='sr-only'>Loading...</span>
              <div className='h-2 w-2 bg-blue-600 rounded-full animate-bounce [animation-delay:-0.3s]'></div>
              <div className='h-2 w-2 bg-blue-600 rounded-full animate-bounce [animation-delay:-0.15s]'></div>
              <div className='h-2 w-2 bg-blue-600 rounded-full animate-bounce'></div>
            </div>
          </div>
        </div>
      )}
      {error && (
        <div className='flex gap-x-8 w-full mx-auto'>
          <div className='min-w-[48px] min-h-[48px]'>
            <Image src='/img/error.png' width={48} height={48} alt='error' />
          </div>
          <div>
            <p className='font-bold'>Ai</p>
            <p className='text-rose-500'>{error}</p>
          </div>
        </div>
      )}
    </div>
  )
}
```

[Link to Frontend Code](https://github.com/vultureprime/ai-web-interface/tree/main/next-chat-interface)

### Backend Development

#### Setting Up a FastAPI Project

Similar to all the previous examples, we choose to use FastAPI as the framework to build an API for our application.

First, install the necessary Python libraries for FastAPI:

```bash
pip install fastapi
pip install "uvicorn[standard]"
```

Create a file named `app.py` and initialize FastAPI:

```python
from fastapi import FastAPI
from fastapi.encoders import jsonable_encoder
from fastapi.responses import JSONResponse
from fastapi.middleware.cors import CORSMiddleware

app = FastAPI()

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)

@app.get("/helloworld")
async def helloworld():
    return {"message": "Hello World"}
```

In the code above, we initialize a FastAPI project and enable CORS for smooth communication with the frontend. To run the server, use the following command:

```bash
uvicorn app:app
```

This command instructs FastAPI to execute the application defined in the `app.py` file, and the server will run on the default port 8000.

Next, create the `data` and `storage` folders to store sample documents and the vector database storage.

#### **Preparation and Ingest Data**

To begin, set up the environment and OpenAI key:

```python
import os
import openai
import dotenv
from llama_hub.file.unstructured.base import UnstructuredReader
from pathlib import Path
from llama_index import VectorStoreIndex, ServiceContext, StorageContext
from llama_index import load_index_from_storage
from llama_index.tools import QueryEngineTool, ToolMetadata
from llama_index.query_engine import SubQuestionQueryEngine
from llama_index.agent import OpenAIAgent
import nest_asyncio
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
from fastapi.responses import JSONResponse

dotenv.load_dotenv()
openai.api_key = os.environ["OPENAI_API_KEY"]
nest_asyncio.apply()
agent = None
```

Continue by loading data into the VectorDB. The example uses data from raw UBER 10-K HTML files for the years 2019-2022:

```python
def read_data(years):
    loader = UnstructuredReader()
    doc_set = {}
    all_docs = []

    for year in years:
        year_docs = loader.load_data(
            file=Path(f"./data/UBER/UBER_{year}.html"), split_documents=False
        )
        # Insert year metadata into each document
        for d in year_docs:
            d.metadata = {"year": year}
        doc_set[year] = year_docs
        all_docs.extend(year_docs)

    return doc_set
```

Now, load the data as documents into the VectorDB, organizing it by year:

```python
def store_data(years, doc_set, service_context):
    index_set = {}

    for year in years:
        storage_context = StorageContext.from_defaults()
        cur_index = VectorStoreIndex.from_documents(
            doc_set[year],
            service_context=service_context,
            storage_context=storage_context,
        )
        index_set[year] = cur_index
        storage_context.persist(persist_dir=f"./storage/{year}")

    return index_set
```

#### Setting Up a Sub Question Query Engine

Create a Query Engine for each year's data by loading the index from the VectorDB:

```python
def load_data(years, service_context):
    index_set = {}

    for year in years:
        storage_context = StorageContext.from_defaults(
            persist_dir=f"./storage/{year}"
        )
        cur_index = load_index_from_storage(
            storage_context, service_context=service_context
        )
        index_set[year] = cur_index

    return index_set
```

Generate Query Engine Tools for each year's data:

```python
def create_individual_query_tool(index_set, years):
    individual_query_engine_tools = [
        QueryEngineTool(
            query_engine=index_set[year].as_query_engine(),
            metadata=ToolMetadata(
                name=f"vector_index_{year}",
                description=f"useful for when you want to answer queries about the {year} SEC 10-K for Uber",
            ),
        )
        for year in years
    ]
    return individual_query_engine_tools
```

#### Synthesize Answers Across the Data

Create a function to synthesize questions for individual query engine tools:

```python
def create_synthesizer(individual_query_engine_tool, service_context):
    query_engine = SubQuestionQueryEngine.from_defaults(
        query_engine_tools=individual_query_engine_tool,
        service_context=service_context,
    )
    return query_engine
```

Generate a Query Engine Tool for sub-question query engine:

```python
def create_sub_question_tool(query_engine):
    query_engine_tool = QueryEngineTool(
        query_engine=query_engine,
        metadata=ToolMetadata(
            name="sub_question_query_engine",
            description="useful for when you want to answer queries that require analyzing multiple SEC 10-K documents for Uber",
        ),
    )
    return query_engine_tool
```

#### Create General Engine

The final engine we are going to create will serve as a query tool used to search for information that is not within the scope of the prepared data. Alternatively, it can be referred to as a chatbot for answering general questions.

```python
def agent_chat():
    chat_engine_tool = [
        QueryEngineTool(
            query_engine=OpenAIAgent.from_tools([]),
            metadata=ToolMetadata(
                name="gpt_agent", description="Agent that can answer general questions."
            ),
        ),
    ]
    return chat_engine_tool
```

#### Create OpenAI Agent from Tools

Combine all query engine tools into an OpenAI Agent:

```python
def build_chat_engine(individual_query_engine_tools, query_engine_tool, gpt_agent):
    global agent
    tools = gpt_agent + individual_query_engine_tools + [query_engine_tool]
    agent = OpenAIAgent.from_tools(tools, verbose=False)
    return agent
```

#### Create RAG API

Build an endpoint for loading data into VectorDB:

```python
@app.post('/buildRAG')
def build_RAG():
    years = [2022, 2021, 2020, 2019]
    service_context = ServiceContext.from_defaults(chunk_size=512)
    doc_set = read_data(years)
    store_data(years, doc_set, service_context)

    return JSONResponse({'status': 'success'})
```

#### Create Chat Endpoint API

Implement an endpoint for processing chat queries:

```python
@app.post('/chat')
def chat(query: str = 'What were some of the biggest risk factors in 2022 for Uber?'):
    global agent
    years = [2022, 2021, 2020, 2019]
    service_context = ServiceContext.from_defaults(chunk_size=512)
    index_set = load_data(years, service_context)
    individual_query_engine_tools = create_individual_query_tool(index_set, years)
    query_engine = create_synthesizer(individual_query_engine_tools, service_context)
    query_engine_tool = create_sub_question_tool(query_engine)
    gpt_agent = agent_chat()

    if agent is None:
        agent = build_chat_engine(individual_query_engine_tools, query_engine_tool, gpt_agent)
    
    answer = agent.chat(query)

    return JSONResponse({'answer': str(answer)})
```

#### Create Utility Endpoint

Implement two additional APIs for utility purposes:

1. Reset the Agent data:

   ```python
   @app.post('/resetChat')
   def resetChat():
       global agent
       agent.reset()
       return JSONResponse({'status': 'complete'})
   ```
2. Use a chatbot without RAG:

   ```python
   @app.post('/chatWithoutRAG')
   def chatWithoutRAG(query: str = 'What were some of the biggest risk factors in 2022 for Uber?'):
       gpt_agent = agent_chat()
       agent = OpenAIAgent.from_tools(gpt_agent, verbose=False)
       answer = agent.chat(query)

       return JSONResponse({'answer': str(answer)})
   ```

#### Deploying and Monitoring on EC2

To deploy FastAPI on an EC2 instance:

1. Create a session using `screen`:

   ```bash
   screen -S name
   ```
2. Navigate to the API folder:

   ```bash
   cd path/to/api
   ```
3. Start FastAPI on port 8000:

   ```bash
   uvicorn app:app --host 0.0.0.0
   ```
4. Detach from the screen session:

   ```bash
   Ctrl+a d
   ```

Now, the server runs in the background even after exiting the session.

#### Setting Up API Gateway and Implementing CORS

To integrate with AWS API Gateway for better management and usage control, add CORS configuration to FastAPI:

```python
from fastapi.middleware.cors import CORSMiddleware

app.app = FastAPI()

app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_credentials=True,
    allow_methods=["*"],
    allow_headers=["*"],
)
```

Connect this server to AWS API Gateway to handle authentication and usage management.

[Link to Backend Code](https://github.com/vultureprime/ai-web-backend/tree/main/Chatbot-openai-llamindex)


# The Beginner's LLM Development Journey

## Journey Overview

As a developer first venturing into the world of Large Language Models (LLMs), you're likely excited to learn how to create projects using this technology. This guide will introduce you to the essential knowledge you need before and during your project development.

We'll provide an overview of the LLM knowledge journey, helping you become a proficient LLM Developer.

{% hint style="info" %}
Welcome to the first version of our guide, released in August 2024. This is just an overview to get you started. We'll be adding more details and examples to each section soon. Keep checking back for updates that will help you on your LLM development journey!
{% endhint %}

***

## Stage 1 : Exploration

The first step of your journey involves exploring familiar concepts.

### Objective

Learn about

* Glossary
* How LLM is work ?&#x20;
* How LLM API is work ?
* LLM behaviour
* Prompt and Completion

### LLM APIs

To begin using LLMs, you should first explore the models available. There are numerous options, but starting with popular models like GPT-4o (from OpenAI) is recommended. Alternatively, you can use our "LLM as a service" offering, which you can learn more about through our resources.

{% content-ref url="/pages/NW8asvdehnm5K81egcWI" %}
[LLM as a service](/getting-started/llm-as-a-service)
{% endcontent-ref %}

{% hint style="info" %}
Learn how to get OpenAI and Anthropic API key from [this](https://blog.float16.cloud/how-to-get-float16-openai-anthropic-api-key/).
{% endhint %}

Once you've obtained your API key, test it using the example commands provided by the LLM provider to ensure you receive a response.

### LLM Behaviour

{% content-ref url="/pages/O5V4k7xPMnmP0ys4POuN" %}
[Variable](/prompting/variable)
{% endcontent-ref %}

{% content-ref url="/pages/7QSRzcZADiM3mn44SGzF" %}
[Condition](/prompting/condition)
{% endcontent-ref %}

{% content-ref url="/pages/xqnKCB7yWfHcIBTbLu4F" %}
[Loop](/prompting/loop)
{% endcontent-ref %}

### Project Example

The fastest way to learn project setup and usage is by studying existing projects or cookbooks from LLM providers. This approach will help you better understand the capabilities of LLMs.

***

## Stage 2 : In-Depth Study

After gaining a basic understanding of how to use LLMs, it's time to go deeper into LLM-specific knowledge

### Objective

Learn about

* Prompt structure
* Prompt technique
* Recursive prompt
* Framework

### Prompt Engineering

Mastering prompt engineering is crucial, as the performance and quality of LLM outputs heavily depend on the prompts you provide. Begin with basic prompting techniques, such as standard usage, role assignment, and task definition. Then, progress to advanced prompting methods like Chain of Thought (CoT) and Few-shot learning.

### Prompt technique

{% content-ref url="/pages/wzN6NHKJgViVJhJ3vNBn" %}
[Demonstration](/prompting/demonstration)
{% endcontent-ref %}

{% content-ref url="/pages/LFle2Np42ZvGydSSwYq4" %}
[Formatting](/prompting/formatting)
{% endcontent-ref %}

{% content-ref url="/pages/G1GcFDmU96vzyBLGoItV" %}
[Chat](/prompting/chat)
{% endcontent-ref %}

{% content-ref url="/pages/XdcS9TntnONQPuLxvoJR" %}
[Technical term (Retrieve)](/prompting/technical-term-retrieve)
{% endcontent-ref %}

### Retrieval Augmented Generation (RAG)

<figure><img src="/files/FPLSvGVi8IGYcVfGUuzq" alt=""><figcaption><p>Retrieval Augmented Generation (RAG) Process</p></figcaption></figure>

RAG is an essential technique for improving LLM performance. It addresses the limitations of pre-trained LLMs, which may have outdated knowledge or lack domain-specific information. RAG allows you to augment the LLM's knowledge base with custom data.

We offer a [RAG example](https://blog.float16.cloud/thai-rag-with-llamaindex-weaviate-seallm/) focusing on Thai language implementation, which you can explore to learn more about this technique.

### Frameworks

Implementing RAG from scratch can be time-consuming. To expedite development and simplify the process, consider using frameworks like LlamaIndex. When working with these frameworks, focus on text processing and data ingestion pipelines, as these are critical components of effective RAG implementation.

By mastering these areas, you'll significantly enhance your ability to create high-performing LLM applications.

***

## Stage 3 : Development

### Objective

Learn about

* How to integrate LLM with other system
* How to define the performance metrics
* How to debugging LLM application

### Applications

It's time to start developing your project. Develop a Simple LLM Project, applying the techniques learned in previous stages to enhance accuracy.

* Apply RAG techniques to your specific data or database. Evaluate whether it's working effectively or if adjustments are needed.
* Conduct a simple evaluation of the results.&#x20;
* Explore advanced RAG methodologies or other techniques to further improve your LLM's accuracy.
* Deploy your RAG system in your personal environment. This step helps you understand the practical aspects of running an LLM application.

By following this development stage, you'll gain practical experience in building and refining LLM applications, setting a strong foundation for more complex projects in the future.

{% hint style="info" %}
The stages beyond this point are tailored for individuals aiming to deploy their LLM projects for public use or within corporate environments.
{% endhint %}

***

## Stage 4 : Proof of Concept (POC)

### Objective

Learn about

* Business use case
* Monitor and Log system
* How to measure the succeed of the POC

### Project

At this stage, you're ready to approach your LLM project more seriously, especially if working with a team. Your goal is to create a compelling proof of concept that demonstrates the project's viability and potential value.

In addition to developing your project with consideration for the techniques learned in previous stages, there are additional tasks you need to focus on:

* Find and document specific business use case scenarios where your LLM application can provide significant value. This evidence will help in gaining buy-in from stakeholders.
* Implement accuracy evaluation protocols to formally assess the accuracy of your LLM application.
* Based on evaluation results, refine and improve your system to meet or exceed accuracy expectations.
* Explore and implement observability and traceability methods for monitoring and tracking the performance of your RAG system. This includes logging, performance metrics, and tools for debugging and analyzing the system's behavior.

By completing this stage, you'll have a robust proof of concept that demonstrates the potential of your LLM application in a real-world business context, setting the stage for potential wider adoption or further development.

***

## Stage 5 : Production

### Objective

Learn about

* How to design scalability system
* SLA and norm

### Deployment

This stage is crucial if you aim to deploy your LLM application with scalability in mind. After successfully completing your Proof of Concept, you'll need to address several additional tasks to prepare for production.

* RAG system optimization, improve the speed of your RAG system to achieve response times of less than 5 seconds per question (request).
* Conduct thorough UAT to ensure your system meets all requirements and user expectations. Once approved, deploy your RAG system to the production environment.

Post-Production Tasks:

* Implement robust monitoring systems, establish appropriate guardrails, and ensure your application complies with relevant regulations and industry standards.
* Learn about scaling your system to handle increased loads. Understand the cost implications of scaling and explore strategies to reduce operational costs.
* Apply your knowledge to scale your RAG system effectively.
* Develop and implement strategies to monetize your LLM application. This may include subscription models, pay-per-use systems, or other revenue-generating approaches.

***

## Congratulations !

You've now completed your journey to creating a potential LLM application. Throughout this project, you've undoubtedly gained valuable knowledge and experience. However, if you have any further questions or topics you'd like to discuss, we're here to help. Please don't hesitate to [contact us](https://discord.gg/j2DVTMjr67) (discord) anytime - we'll be glad to assist you and will respond as quickly as possible!

This journey is just the beginning, and we look forward to hearing about your continued success and innovations in the world of LLM applications.


# Glossary

Essential terminology in the world of Large Language Models (LLM)

## Welcome to our LLM Glossary!

This comprehensive resource is designed for newcomers to the world of Large Language Models (LLMs). Here, you'll find explanations for terms that might leave you puzzled at first encounter.

Our glossary compiles key terminology crucial for understanding LLMs, tailored for software developers and AI enthusiasts alike. Each entry offers a concise, easy-to-understand explanation, arranged alphabetically for your convenience.

Our glossary is *multi-language support*. We currently offer definitions in English and Thai, with plans to expand to more languages in the future.

> This feature is particularly beneficial for Thai developers, as we aim to empower them by providing resources in their native language, bridging the gap often created by language barriers. While supporting local developers, we maintain a global outlook, ensuring that knowledge is accessible to everyone worldwide.

Use this glossary as your go-to reference for grasping core concepts and specialized vocabulary in LLM technology.

Ready to explore? Check out our glossary now!

{% content-ref url="/pages/471qCg46Z7eWmwnuH785" %}
[\[English Version\] LLM Glossary](/journey/glossary/english-version-llm-glossary)
{% endcontent-ref %}

Thai developers, click here for the Thai version.

{% content-ref url="/pages/a7YU7tFzPFYgJa7nR8aE" %}
[\[ภาษาไทย\] LLM Glossary](/journey/glossary/llm-glossary)
{% endcontent-ref %}

{% hint style="info" %}
For any feedback or suggestions, please reach out to us on [Discord](https://discord.gg/j2DVTMjr67).
{% endhint %}


# \[English Version] LLM Glossary

LLM Glossary in English Language

<table><thead><tr><th>Table of contents</th><th data-hidden></th><th data-hidden></th><th data-hidden></th></tr></thead><tbody><tr><td><a href="#agents">Agents</a></td><td></td><td></td><td></td></tr><tr><td><a href="#few-shot-learning">Few-shot learning</a></td><td></td><td></td><td></td></tr><tr><td><a href="#fine-tuning">Fine-tuning</a></td><td></td><td></td><td></td></tr><tr><td><a href="#framework">Framework</a></td><td></td><td></td><td></td></tr><tr><td><a href="#guardrail">Guardrail</a></td><td></td><td></td><td></td></tr><tr><td><a href="#hallucination">Hallucination</a></td><td></td><td></td><td></td></tr><tr><td><a href="#large-language-models-llms">LLM</a></td><td></td><td></td><td></td></tr><tr><td><a href="#overfitting">Overfitting</a></td><td></td><td></td><td></td></tr><tr><td><a href="#pre-trained">Pre-trained</a></td><td></td><td></td><td></td></tr><tr><td><a href="#prompt-engineering">Prompt Engineering</a></td><td></td><td></td><td></td></tr><tr><td><a href="#quantization">Quantization</a></td><td></td><td></td><td></td></tr><tr><td><a href="#retrieval-augmented-generation-rag">RAG</a></td><td></td><td></td><td></td></tr><tr><td><a href="#token">Token</a></td><td></td><td></td><td></td></tr><tr><td><a href="#vector-database">Vector Database</a></td><td></td><td></td><td></td></tr></tbody></table>

## Agents&#x20;

An Agent is a subset of an LLM (Large Language Model) within a larger system, possessing the same capabilities as any LLM (i.e., understanding human language). However, it is programmed to execute specific tasks as instructed. Agents process and execute tasks under these instructions. The separation of Agents allows for more focused and efficient performance tailored to specific objectives.

For example, an LLM Chatbot providing customer service might have primary services such as order processing and answering frequently asked questions (FAQ). We can simply divide the Agents into three types:

Manager Agent: Welcomes customers and delegates tasks to the relevant Agents. Order Agent: Responsible for processing orders. FAQ Agent: Responsible for answering frequently asked questions. For instance, when a customer first engages with the system, they will encounter the Manager Agent. If the customer wishes to place an order, the Manager Agent processes this request and delegates it to the Order Agent, which specializes in handling orders.

<figure><img src="/files/E0LsHvon16q3bVSfPpiK" alt=""><figcaption><p>Example of agent usage</p></figcaption></figure>

You can learn more about LLM Agents at this [Link](https://www.vultureprime.com/blogs/what-is-llm-agents).

## Few-shot learning

Few-shot learning is a method by which an LLM (Large Language Model) can be trained with only a few examples. In the context of LLMs, if we want to teach the model to translate simple regional dialects, we can start by teaching it that the word "บ่" means "ไม่" (no/not) and provide an example. When faced with a question requiring translation of the regional dialect, the LLM can then translate dialects containing the word "บ่."

You can see additional examples at the link below.

{% embed url="<https://prompt.float16.cloud/prompt/8663be85-7c85-48ed-b984-81651069f55e>" %}

## Fine-tuning

Fine-tuning is the process of adjusting a pre-trained LLM (Large Language Model) or a model trained from scratch to better suit the specific tasks we intend to use it for. This is accomplished by training the model with additional or specialized data, adjusting certain parameters (such as temperature), or managing prompts appropriately.

For example, if you want to use an LLM in the medical field, you may need to fine-tune the LLM to ensure the quality of responses is suitable for medical use. This is because most base LLMs are not configured to respond in a medical context or contain extensive medical information.

You can learn more about using prompts with technical terms from the link below.

{% content-ref url="/pages/XdcS9TntnONQPuLxvoJR" %}
[Technical term (Retrieve)](/prompting/technical-term-retrieve)
{% endcontent-ref %}

## Framework

A framework or data framework is a structural model that facilitates the connection of processes from start to finish, enabling the creation of LLM applications more efficiently and easily. You can use the functions of a framework to develop LLM applications without starting from scratch, which significantly reduces development time.

Just as JavaScript has various frameworks like Angular, Vue.js, or React, there are several frameworks available for developing LLM applications. Popular ones include LlamaIndex and Langchain.

You can study an example of using LlamaIndex as a framework for creating a chatbot from the link below.

{% embed url="<https://blog.float16.cloud/thai-rag-with-llamaindex-weaviate-seallm/>" %}

## Guardrail

Guardrail is a technique used to set boundaries for the responses of an LLM (Large Language Model). By defining rules or constraints, you can control the LLM's operations, preventing errors or deviations from its intended purpose.

For example, if you create an LLM Chatbot for hospital services, you can use the Guardrail technique to ensure that the Chatbot only responds to topics related to your hospital. Without these constraints, users might use the Chatbot for unrelated tasks, potentially causing damage to your business.

You can learn more about applying Guardrail to Chatbots from this [Link](https://www.vultureprime.com/blogs/openai-with-guardrail).

## Hallucination

Hallucination is an event where an LLM (Large Language Model) or NLP (Natural Language Processing) system generates information or text that does not align with reality or lacks supporting data, resulting in incorrect outputs.&#x20;

This can occur due to various reasons, such as incomplete or noisy training data, ambiguous questions, or model bias from the learned data. Hallucinations can significantly impact applications requiring high accuracy, such as in medicine, law, or engineering.

<figure><img src="/files/tw1jiQAPurGcTAKdjjkC" alt=""><figcaption><p>ChatGPT's Hallucination</p></figcaption></figure>

## Large Language Models (LLMs)

A large language model is a sizable AI model trained to understand and generate human language. It acts like an artificial brain that consumes vast amounts of linguistic data, learning to comprehend meanings, contexts, and various ways of using language. LLMs are developed based on deep learning algorithms and serve as the foundation for generative AI.

Examples of LLMs include GPT-4, Claude 3.5 Sonnet, and SeaLLM v3. Some providers offer playgrounds and APIs for developers to experiment with their models.

For those interested in trying out SeaLLM through LLM as a Service, you can use the App Float16.cloud via the link below.

{% embed url="<https://app.float16.cloud>" %}

## Overfitting

Overfitting is a behavior of an AI model where it learns the details of a specific training dataset too well. Besides learning useful information, the model may also learn noise in the data. This can happen when the model is too complex. When predicting data it has seen before, the model can be very accurate, but when it encounters new, unseen data, it may fail to make correct predictions and perform poorly.

For example, suppose we create a model to predict whether an image contains kitchen utensils. If the training images all have a gas stove in the background, the model might learn to associate kitchen utensils with the presence of a gas stove. If it then encounters an image of kitchen utensils without a gas stove in the background, it might incorrectly predict that the image does not contain kitchen utensils.

In the context of LLMs, overfitting can occur when we create a model to help write stories. The model might produce responses very similar to the training text, lacking creativity. This is problematic for models involved in creative tasks or generative AI.

In addition to overfitting, there is also underfitting, which occurs when an AI model has too little training data, leading to incorrect answers and high bias.

## Pre-trained

Pre-trained models are AI models that have already been trained on large datasets. These models are typically trained with diverse and comprehensive data to understand and handle various languages or information effectively.

Training large models from scratch requires significant resources and time. However, using pre-trained models can reduce the time and resources needed to train a new model. These pre-trained models can be used and further improved (fine-tuned) to suit specific tasks more easily.

Examples of pre-trained language models include Llama3, Gemma, and Mamba. These models can be found on platforms like [Hugging Face](https://huggingface.co/models)

## Prompt Engineering&#x20;

A "prompt" is a text or question that we input into an LLM (Large Language Model) to instruct it to generate information and answers. However, the responses may not always align with our expectations. Therefore, we need to engage in a process called "Prompt Engineering" to ensure that the answers produced meet our requirements.

"Prompt Engineering" can occur during the creation of the prompt or while asking questions. This involves specifying conditions more clearly or configuring the LLM to have a particular character or response style as we desire.

You can learn about Basic Prompting from the link below.

{% content-ref url="/pages/j0mfSRaz5h0fk5HqNs2E" %}
[Prompting](/prompting/variable)
{% endcontent-ref %}

## Quantization&#x20;

Quantization is the process of reducing the precision of the parameters in an AI model to make the model smaller, save memory, and increase computational speed, without significantly compromising performance.

You can view the Benchmark LLM Speed Inference from Float16 at the link below.

{% embed url="<https://quantize.float16.cloud/>" %}

## Retrieval Augmented Generation (RAG)

Retrieval Augmented Generation is a technique used to enhance the performance of LLMs (Large Language Models) by providing them with specialized or updated information, enabling the models to answer more specific questions accurately.

Since pre-trained models are trained with data limited to specific time periods or broad information, we need to provide the models with the specific data we want them to know. However, direct training requires substantial resources. Therefore, the RAG process is a popular technique, which involves the following steps:

<figure><img src="/files/YKOSd5Z1GhPEpk36c9wo" alt=""><figcaption></figcaption></figure>

1. Embed the additional data into vectors and store them in a vector database.
2. Input the question into the LLM.
3. The LLM converts the question into a vector and searches for relevant information in the vector database.
4. The LLM processes the relevant information into human language and responds.

You can learn more about RAG from this [Link](https://www.vultureprime.com/blogs/rag-101-step-by-step-introduction).

## Token&#x20;

A token is the smallest unit of data that is segmented for processing. It can be a word, character, or other subunit used in data analysis and processing. Each model and language may use different methods to count tokens, which can be in the form of words, syllables, or characters.

Most LLMs use tokens to indicate processing speed, typically measured in tokens per second (Token/Second), representing the number of tokens that can be processed or generated per second. The higher the number, the faster the response time. Additionally, the pricing model for LLM as a Service is often based on the cost per 1 million tokens.

You can try out the Tokenizer Playground at the link below.

{% embed url="<https://float16.cloud/tokenizer>" %}
Tokenizer
{% endembed %}

## Vector Database

A vector database is a type of database designed to store and search data in the form of vectors. Generally, vectors are sets of numbers that represent features or various types of data, such as image data, text data, or audio data. Using vectors to represent data makes it easier to perform calculations or statistical analyses.

Vector databases are particularly important in the context of LLMs (Large Language Models), especially for similarity searches. These include tasks such as finding similar images, searching for text with similar meanings, or finding similar audio clips.

You can learn more about Vector Databases from the link below.

{% embed url="<https://www.vultureprime.com/blogs/vector-database>" %}


# \[ภาษาไทย] LLM Glossary

LLM Glossary ฉบับภาษาไทย

<table><thead><tr><th>สารบัญ</th><th data-hidden></th><th data-hidden></th><th data-hidden></th></tr></thead><tbody><tr><td><a href="#agents">Agents</a></td><td></td><td></td><td></td></tr><tr><td><a href="#few-shot-learning">Few-shot learning</a></td><td></td><td></td><td></td></tr><tr><td><a href="#fine-tuning">Fine-tuning</a></td><td></td><td></td><td></td></tr><tr><td><a href="#framework">Framework</a></td><td></td><td></td><td></td></tr><tr><td><a href="#guardrail">Guardrail</a></td><td></td><td></td><td></td></tr><tr><td><a href="#hallucination">Hallucination</a></td><td></td><td></td><td></td></tr><tr><td><a href="#large-language-models-llms">LLM</a></td><td></td><td></td><td></td></tr><tr><td><a href="#overfitting">Overfitting</a></td><td></td><td></td><td></td></tr><tr><td><a href="#pre-trained">Pre-trained</a></td><td></td><td></td><td></td></tr><tr><td><a href="#prompt-engineering">Prompt Engineering</a></td><td></td><td></td><td></td></tr><tr><td><a href="#quantization">Quantization</a></td><td></td><td></td><td></td></tr><tr><td><a href="#retrieval-augmented-generation-rag">RAG</a></td><td></td><td></td><td></td></tr><tr><td><a href="#token">Token</a></td><td></td><td></td><td></td></tr><tr><td><a href="#vector-database">Vector Database</a></td><td></td><td></td><td></td></tr></tbody></table>

## Agents&#x20;

คือ LLM ย่อยในระบบใหญ่ที่มีความสามารถเหมือน LLM ทุกประการ (กล่าวคือ ความเข้าใจภาษามนุษย์) แต่จะมีคำสั่งให้ดำเนินการบางอย่างตามที่กำหนด Agents จะทำการประมวลผลและดำเนินการภายใต้คำสั่งนั้น ซึ่งการแยก Agents ช่วยให้สามารถแยกการทำงานและเพิ่มประสิทธิภาพ Agent ได้ตรงตามจุดประสงค์การทำงานมากขึ้น

ยกตัวอย่างเช่น LLM Chatbot ที่ให้บริการลูกค้า อาจมีบริการหลักๆ คือการสั่งซื้อสินค้าและการตอบคำถามที่พบบ่อย (FAQ) เราสามารถแบ่ง Agents แบบง่ายๆ ได้เป็น 3 ตัว ได้แก่

* Manager Agent สำหรับต้อนรับลูกค้าและส่ง Task งานไปยังแต่ละ Agents ที่เกี่ยวข้อง
* Order Agent มีหน้าที่สำหรับสั่งซื้อสินค้า
* FAQ Agent มีหน้าที่สำหรับตอบคำถามที่พบบ่อย

เช่น เมื่อลูกค้าเข้ามาครั้งแรกจะพบกับ Manager Agent ก่อน และเมื่อลูกค้าต้องการสั่งสินค้า Manager Agent จะทำการประมวลผลและส่งงานไปยัง Order Agent ที่มีความสามารถในการสั่งซื้อสินค้าต่อไป

<figure><img src="/files/E0LsHvon16q3bVSfPpiK" alt=""><figcaption><p>ตัวอย่างการใช้งาน Agent</p></figcaption></figure>

สามารถศึกษาเพิ่มเติมเกี่ยวกับ LLM Agents ได้ตาม [Link](https://www.vultureprime.com/blogs/what-is-llm-agents) นี้

## Few-shot learning

คือวิธีการเรียนรู้ของ LLM ที่สามารถฝึกฝนโมเดลด้วยตัวอย่างเพียงไม่กี่ตัวอย่าง ตัวอย่างในบริบทของ LLM เช่น หากเราต้องการสอนให้ LLM แปลภาษาถิ่นคำง่ายๆ เราเริ่มจากการสอนคำว่า "บ่" ที่แปลว่า "ไม่" พร้อมยกตัวอย่าง เมื่อเจอคำถามที่ต้องการให้แปลภาษาถิ่น LLM ก็จะสามารถแปลภาษาถิ่นที่มีคำว่า "บ่" ได้

สามารถดูตัวอย่างเพิ่มเติมได้ตามลิงก์ด้านล่าง

{% embed url="<https://prompt.float16.cloud/prompt/8663be85-7c85-48ed-b984-81651069f55e>" %}

## Fine-tuning

คือ กระบวนการปรับแต่ง LLM ที่ผ่านการฝึกฝนมาแล้ว ไม่ว่าจะเป็น Pre-trained Model หรือ From Scratch Model ให้เข้ากับงานที่เราต้องการนำไปใช้ ด้วยการฝึกฝนด้วยข้อมูลเพิ่มเติมหรือข้อมูลเฉพาะทาง การปรับค่าบางอย่าง (เช่น Temperature) หรือ การจัดการ Prompt ให้เหมาะสม&#x20;

เช่น ถ้าหากคุณต้องการใช้ LLM ในงานด้านการแพทย์ คุณอาจต้องมีการ Fine-tuning LLM เพื่อให้คุณภาพของคำตอบสามารถนำไปใช้ในทางการแพทย์ได้จริง เพราะ LLM พื้นฐานส่วนใหญ่ จะไม่ได้ตั้งค่าให้ตอบในเชิงการแพทย์หรือมีข้อมูลทางการแพทย์มากนักf

สามารถศึกษาเรื่องการใช้ Prompt แบบ Technical Term ได้จากลิงก์ด้านล่าง

{% content-ref url="/pages/XdcS9TntnONQPuLxvoJR" %}
[Technical term (Retrieve)](/prompting/technical-term-retrieve)
{% endcontent-ref %}

## Framework

หรือ Data Framework เป็นโครงสร้างการทำงานที่ช่วยอำนวยความสะดวกในการเชื่อมต่อกระบวนการทำงานตั้งแต่เริ่มต้นจนจบ ทำให้สามารถสร้าง LLM Application ได้อย่างมีประสิทธิภาพและง่ายดายมากยิ่งขึ้น คุณสามารถใช้ฟังก์ชันของ Framework เพื่อพัฒนา LLM Application โดยไม่ต้องเริ่มต้นจากศูนย์ ซึ่งจะช่วยลดระยะเวลาในการพัฒนาได้มาก

เช่นเดียวกับที่ JavaScript มี Framework หลากหลายตัว เช่น Angular, Vue.js หรือ React ในการพัฒนา LLM Application ก็มี Framework ให้เลือกใช้หลายตัวเช่นกัน ที่นิยมใช้ได้แก่ LlamaIndex และ Langchain

คุณสามารถศึกษาตัวอย่างการใช้ LlamaIndex เป็น Framework ในการสร้าง Chatbot ได้ตามลิงก์ด้านล่าง

{% embed url="<https://www.vultureprime.com/how-to/how-to-create-thai-rag-with-llamaindex-weaviate-seallm>" %}

## Guardrail

เป็นเทคนิคหนึ่งเพื่อกำหนดขอบเขตการให้คำตอบของ LLM โดยสามารถกำหนดกฎหรือข้อจำกัดในการใช้งานเพื่อควบคุมการทำงานของ LLM ป้องกันไม่ให้เกิดข้อผิดพลาดหรือผิดวัตถุประสงค์ในการใช้งาน

เช่น คุณสร้าง LLM Chatbot สำหรับให้บริการในโรงพยาบาล คุณใช้เทคนิค Guardrail เพื่อควบคุมการทำงานของ Chatbot ให้ตอบเฉพาะเรื่องที่เกี่ยวข้องกับโรงพยาบาลของคุณเท่านั้น ซึ่งถ้าคุณไม่เขียนข้อจำกัดไว้ก็อาจมีคนมาใช้งาน Chatbot คุณไปทำงานอย่างอื่น ซึ่งจะก่อให้เกิดความเสียหายแก่ธุรกิจของคุณได้

สามารถศึกษาเรื่องการประยุกต์ใช้ Guardrail กับ Chatbot ได้ตาม [Link](https://www.vultureprime.com/blogs/openai-with-guardrail) นี้

## Hallucination

เป็นเหตุการณ์ที่ LLM หรือ NLP สร้างข้อมูลหรือข้อความที่ไม่ตรงกับความเป็นจริงหรือไม่มีข้อมูลรองรับ ทำให้ผลลัพธ์ที่ได้ไม่ถูกต้อง&#x20;

ซึ่งเกิดได้จากหลายสาเหตุ เช่น Training Data ที่ไม่สมบูรณ์หรือมีข้อมูลขยะเยอะเกินไป คำถามที่กำกวม หรือ Bias ของตัวโมเดลจากการข้อมูลที่เรียนรู้ ซึ่งอาจจะส่งผลกระทบต่อการใช้งานในส่วนงานที่ต้องการความแม่นยำสูง เช่น การแพทย์, กฎหมาย หรือ วิศวกรรม

<figure><img src="/files/tw1jiQAPurGcTAKdjjkC" alt=""><figcaption><p>ตัวอย่างการ Hallucination ของ ChatGPT</p></figcaption></figure>

## Large Language Models (LLMs)

คือโมเดล AI ขนาดใหญ่ที่ถูกฝึกฝนมาให้เข้าใจและสร้างภาษามนุษย์ได้ เป็นเหมือนสมองกลที่กินข้อมูลภาษามหาศาลเข้าไป แล้วเรียนรู้ที่จะเข้าใจความหมาย บริบท และวิธีการใช้ภาษาในรูปแบบต่าง ๆ ถูกพัฒนาต่อยอดมาจากอัลกอริทึม Deep Learning และเป็นโมเดลพื้นฐานของ Generative AI&#x20;

ตัวอย่าง LLMs เช่น GPT4-o, Claude 3.5 Sonnet หรือ SeaLLM v3 โดยบางเจ้าก็จะมี Playground และ API สำหรับให้ developer ไปทดลองใช้งานโมเดลของตัวเองได้

สำหรับคนที่อยากทดลองใช้งาน SeaLLM ผ่าน LLM as a Service สามารถใช้งาน App Float16.cloud ได้ตามลิงก์ด้านล่าง

{% embed url="<https://app.float16.cloud>" %}

## Overfitting

คือพฤติกรรมการเรียนรู้ของโมเดล AI ที่เรียนรู้รายละเอียดของ Train Data ชุดหนึ่งๆมากเกินไป ซึ่งนอกจากข้อมูลดีๆแล้วยังอาจเรียนไปถึงข้อมูล noise ด้วย ซึ่งเกิดได้ในกรณีที่โมเดลมีความซับซ้อน หากทายข้อมูลที่เคยเรียนรู้จะทายได้ถูกต้องแม่นยำมาก แต่เมื่อทำนายข้อมูลที่ไม่เคยรู้จักมาก่อนโมเดลจะตอบคำถามได้ไม่ถูกต้องและไม่สามารถทำงานได้ดีได้ในข้อมูลใหม่

เช่น สมมติเราสร้างโมเดลสำหรับทำนายว่ารูปภาพนี้เป็นเครื่องครัวหรือไม่ แล้วรูปภาพที่เราใช้เทรนเป็นรูปภาพที่มีฉากหลังติดเตาแก๊สหมด หากโมเดลเจอรูปภาพที่เป็นเครื่องครัวแต่ฉากหลังไม่มีเตาแก๊สก็อาจจะตอบว่าไม่ใช่เครื่องครัวก็เป็นได้

หรือใน LLM การเกิด Overfitting เช่น เราสร้างโมเดลสำหรับช่วยเขียนนิทาน ก็อาจจะได้คำตอบที่ใกล้เคียงกับข้อความที่ใช้ฝึกฝนมากๆ ไม่มีความคิดสร้างสรรค์ ซึ่งจำเป็นสำหรับโมเดลที่เกี่ยวข้องกับการสร้างสรรค์สิ่งใหม่ๆ หรือ Generative AI

ซึ่งนอกจาก Overfitting แล้ว ยังมี Underfitting ซึ่งก็คือการที่โมเดล AI มี Train Data น้อยเกินไป ทำให้ตอบคำถามไม่ถูก และมีความ Bias สูง

## Pre-trained

คือโมเดลที่ผ่านการฝึกฝนด้วยข้อมูลขนาดใหญ่มาแล้ว โดยโมเดลเหล่านี้มักจะถูกฝึกด้วยข้อมูลที่หลากหลายและครอบคลุม เพื่อให้สามารถเข้าใจและจัดการกับภาษาหรือข้อมูลที่หลากหลายได้อย่างมีประสิทธิภาพ

การฝึกโมเดลขนาดใหญ่จากศูนย์ (scratch) ต้องใช้ทรัพยากรและเวลาอย่างมาก แต่การใช้โมเดล pre-trained จะช่วยลดเวลาและทรัพยากรที่ต้องใช้ในการฝึกโมเดลใหม่ ทำให้สามารถนำโมเดลที่ผ่านการฝึกมาแล้วมาใช้และปรับปรุงเพิ่มเติม (fine-tuning) เพื่อให้เหมาะสมกับงานเฉพาะทางได้ง่ายขึ้น

ตัวอย่าง Pre-trained Model ทางภาษา เช่น Llama3, Gemma, Mamba เป็นต้น สามารถค้นหาโมเดลได้ใน [Hugging Face](https://huggingface.co/models)

## Prompt Engineering&#x20;

"Prompt" คือข้อความหรือคำถามที่เราป้อนให้ LLM ซึ่งทำหน้าที่เป็นคำสั่งให้ LLM นำข้อมูลและคำตอบออกมาให้ แต่คำตอบที่ออกมาอาจไม่ตรงกับที่เราต้องการเสมอไป ดังนั้นเราจำเป็นต้องทำกระบวนการ "Prompt Engineering" เพื่อให้คำตอบที่ได้มาเป็นผลลัพธ์ที่ตรงกับความต้องการของเรา

การทำ "Prompt Engineering" อาจเกิดขึ้นระหว่างการสร้าง Prompt หรือขณะถามคำถาม โดยการกำหนดเงื่อนไขต่างๆ ให้ชัดเจนมากขึ้น หรือการตั้งค่า LLM ให้มีคาแรคเตอร์หรือลักษณะการตอบคำถามตามที่เรากำหนด&#x20;

คุณสามารถศึกษา Basic Prompting ได้ตามลิงก์ด้านล่าง

{% content-ref url="/pages/j0mfSRaz5h0fk5HqNs2E" %}
[Prompting](/prompting/variable)
{% endcontent-ref %}

## Quantization&#x20;

คือกระบวนการลดความละเอียดของค่าพารามิเตอร์ในโมเดล AI เพื่อให้โมเดลมีขนาดเล็กลง ประหยัดหน่วยความจำและทำงานได้เร็วขึ้น โดยไม่สูญเสียประสิทธิภาพมากนัก

สามารถดู Benchmark LLM Speed Inference จาก Float16 ได้ที่ลิงก์ด้านล่าง

{% embed url="<https://quantize.float16.cloud/>" %}

## Retrieval Augmented Generation (RAG)

เป็นเทคนิคที่ใช้ในการเพิ่มประสิทธิภาพของ LLM โดยการป้อนข้อมูลเฉพาะทางหรืออัปเดตข้อมูลให้กับโมเดล ทำให้โมเดลสามารถตอบคำถามที่เฉพาะทางมากขึ้นได้

เนื่องจาก Pre-trained Model ต่างๆ ถูกเทรนด้วยข้อมูลจำกัดช่วงเวลา หรือ ข้อมูลกว้างๆ ทำให้เราต้องนำข้อมูลที่เราต้องการให้โมเดลรู้มาฝึกฝนเพิ่มเอง แต่เราจะไม่ได้ทำการฝึกฝนตรงๆ เพราะการฝึกฝนโมเดลจำเป็นต้องใช้ทรัพยากรจำนวนมาก ดังนั้นกระบวนการ RAG จึงเป็นเทคนิคที่นิยมทำกัน ซึ่งมีขั้นตอนคร่าวๆ ดังนี้

<figure><img src="/files/ZF4EcvE7K0COUESILzpO" alt=""><figcaption></figcaption></figure>

1. นำข้อมูลเพิ่มเติมไป Embedding แปลงเป็น Vector และเก็บไว้ใน Vector Database
2. ป้อนคำถามกับ LLM
3. LLM แปลงคำถาม เป็น Vector และนำไปค้นหาข้อมูลที่เกี่ยวข้องใน Vector Database&#x20;
4. LLM นำข้อมูลที่เกี่ยวข้องมาประมวลผลเป็นภาษามนุษย์ และ Response กลับไป

โดยสามารถศึกษาเกี่ยวกับ RAG เพิ่มเติมได้ตาม [Link](https://www.vultureprime.com/blogs/rag-101-step-by-step-introduction) นี้

## Token&#x20;

คือหน่วยย่อยที่สุดของข้อมูลที่ถูกแบ่งออกมาเพื่อการประมวลผล อาจเป็นคำ (words), อักขระ (characters), หรือหน่วยย่อยอื่นๆ ที่ใช้ในการวิเคราะห์และประมวลผลข้อมูล ซึ่งแต่ละโมเดล แต่ละภาษาจะใช้วิธีนับจำนวน Token ต่างกัน เป็นคำ พยางค์ หรืออักขระก็ได้

ส่วนใหญ่แล้ว LLM จะใช้ Token ในการบ่งบอกความเร็วในการทำงานเช่นกัน โดยจะเป็นหน่วย Token/Second หรือจำนวน Token ที่สามารถรับ หรือ ส่งออก ต่อวินาที ยิ่งมากก็แสดงว่าความเร็วในการตอบสนองก็จะเร็วมาก หรือ Pricing Model ของ LLM as a Service ก็จะคิดเป็นราคา ต่อ 1 Million Token เช่นกัน

สามารถทดลองใช้ Tokenizer Playground ได้ตามลิงก์ด้านล่าง

{% embed url="<https://float16.cloud/tokenizer>" %}
Tokenizer
{% endembed %}

## Vector Database

คือประเภทของ Database ชนิดหนึ่ง ที่ออกแบบมาเพื่อจัดเก็บและค้นหาข้อมูลในรูปแบบของเวกเตอร์ (vector) โดยทั่วไปแล้ว เวกเตอร์จะเป็นชุดของตัวเลขที่แทนค่าคุณลักษณะหรือข้อมูลต่าง ๆ เช่น ข้อมูลภาพ ข้อมูลข้อความ หรือข้อมูลเสียง ซึ่งการใช้เวกเตอร์ในการแทนค่าข้อมูลทำให้สามารถนำไปใช้ในงานที่ต้องการการคำนวณหรือการวิเคราะห์เชิงสถิติได้ง่ายขึ้น

Vector database มีความสำคัญในงาน LLM โดยเฉพาะกับการค้นหาที่มีความคล้ายคลึงกัน (Similarity) เช่น การค้นหาภาพที่คล้ายกัน การค้นหาข้อความที่มีความหมายคล้ายกัน หรือการค้นหาเสียงที่คล้ายกัน

ศึกษาเกี่ยวกับ Vector Database เพิ่มเติมได้ตามลิงก์ด้านล่าง

{% embed url="<https://www.vultureprime.com/blogs/vector-database>" %}


# How to install node

node installation guide

This guide will help you install Node.js using Node Version Manager (NVM). NVM allows you to easily install and manage different versions of Node.js on your system.

## Installation Steps

### 1. Install NVM (Node Version Manager)

Open your terminal and run the following command:

```bash
curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
```

After installation, add these lines to your shell configuration file (`.bashrc`, `.zshrc`, or `.profile`):

```bash
export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh"
[ -s "$NVM_DIR/bash_completion" ] && \. "$NVM_DIR/bash_completion"
```

**Note**: You may need to restart your terminal or run `source ~/.bashrc` (or your appropriate shell configuration file) for the changes to take effect.

### 2. Verify NVM Installation

Verify that NVM is properly installed by checking its version:

```bash
nvm --version
```

### 3. View Available Node.js Versions

To see all available Node.js versions:

```bash
nvm ls-remote
```

This command displays a list of all Node.js versions that can be installed through NVM.

### 4. Install Node.js

Install your desired version of Node.js (version 20 or higher is required for float16 CLI):

```bash
nvm install <node_version>
```

### 5. Verify Installation

After installation, verify that Node.js is properly installed:

```bash
node --version
```


# Variable

### Intro

LLMs have the ability to recall and reference the data we declared.

We could compare this ability to declaring variables in any programming language.

### How it work ?&#x20;

You could use "ANY" method to declare a variable.

```
i.e.
a = "hello world"
a => "hello world"
a -> "hello world"
<a>helloworld<a>
{a}helloworld{a}
a > "hello world"
a => hello world
"""a
hello world
"""
```

{% hint style="info" %}
A software developer could consider an LLM as a 'natural coding language.' It has similar abilities to other programming languages.
{% endhint %}

### Prompt Example

{% embed url="<https://prompt.float16.cloud/prompt/4039d295-c037-4c7b-999d-0ddb2a07663f>" %}
Basic declaration and printing of a variable's value
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/71382ab9-b152-4b54-a66c-0fb7ecfae1ed>" %}
Basic declaration and printing of a variable's value using some variable
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/ccb791c7-6c98-4ee6-86ec-7dd3c5b4594d>" %}
Change value of variable
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/9836b258-ef9a-474f-b383-614d8ce08647>" %}
Access nested variable
{% endembed %}


# Condition

### Intro

LLM have the ability to make decisions and take action if the "target value" matches certain criteria.&#x20;

The rules can be determined by numbers, strings, or **real-world characteristics**.

### How it work ?

LLM can understand not only text but also have the ability to comprehend common sense. This is because LLM are trained on HUGE datasets that include common sense knowledge.

```
word list => [Hello, Sandwich, Food, Bob]
if word list is type of food change into Food.
---------------------------------------------------
<Email> Hello Nathan, 
How about you guy ? 
Hope you get well soon.
Best 
Bob.
<Email>

Remove name in <Email> and replace with placeholder.
```

We can prompt an LLM to make decisions and take action by setting criteria for the decision and defining the next action when the criteria are met.

{% hint style="info" %}
A software developer could consider this ability like if-else ability to other programing languages.
{% endhint %}

### Prompt example

{% embed url="<https://prompt.float16.cloud/prompt/d633841d-0972-4103-b9b4-94981956aa88>" %}
Replace some word into anther word
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/34ca39f0-c313-4bc7-8a34-4a691597ee71>" %}
Replace some word into color
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/6fd7f373-ca22-4096-9bc2-d134d9640293>" %}
Replace some word and group
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/7293d439-d7db-4da8-bc15-795dca1686e0>" %}
Get some word with emotion condition
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/6ef67164-7ed6-4cd4-a5d3-48fd9cd26161>" %}
Remove name and replace with placeholder
{% endembed %}


# Demonstration

### Intro

LLMs have the ability to follow demonstrations by example that we have provided, without any explicit instructions or conditions. It's like providing X as input and Y as output.&#x20;

If we have enough pairs to demonstrate, the LLM will learn and try to predict Y for incoming X.

### How it work ?

LLMs have an ability called "in-context learning."

In-context learning comes along with LLMs by providing examples called few-shot or many-shot examples in the prompt.&#x20;

If you have some experience with training machine learning models, you will notice this concept is similar to preparing a dataset for supervised learning by providing x\_train and y\_train.

```
Rice => noun
Eat => verb
Sleep => verb
Food => 
```

{% hint style="info" %}
A software developer could consider this ability to be a **new type** of capability for programming languages. In traditional programming languages, you need to define conditions to modify the input. However, with this ability, you just prepare example pairs of input and output for the LLM. You no longer need to write conditions to modify the input.
{% endhint %}

### Prompt example

{% embed url="<https://prompt.float16.cloud/prompt/240fbb48-5175-471c-ac72-45012b2e386e>" %}
Machine translation
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/af7f1030-c94e-4787-91fc-a80f8d6aa23a>" %}
Part of speech
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/7911b435-0298-488d-82b6-70b0965318b3>" %}
Arrange text by demonstration
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/ee4a228f-9d45-4f24-abe7-454a287548dd>" %}
Information extraction
{% endembed %}


# Loop

### Intro

LLMs have the ability to loop based on a modifier or condition to determine the action.

### How it work ?

LLMs can operate step by step. The number of loops can be defined by a specific number or can loop through the entire data.

```
food list = []

Add food name into food list 10 times.
```

{% hint style="info" %}
A software developer could consider this ability similar to the **Loop** functionality in other programming languages.
{% endhint %}

### Prompt example

{% embed url="<https://prompt.float16.cloud/prompt/e8c57ab1-73ce-4a7c-a816-9a69f239269f>" %}
Iterative add name of food
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/d1d2fe39-f3fc-4a0b-96b1-aebb75627dce>" %}
Iterative loop with condition
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/5832987b-1b91-4737-8b99-56aaab5a2a24>" %}
Iterative loop and stop (While loop)
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/51ee2c08-048c-43cd-af59-9cb9d3527fcc>" %}
Iterative loop and stop (For loop)
{% endembed %}


# Formatting

### Intro

Text formatting is a crucial ability of LLMs because LLMs can understand the text that we provide to them. LLMs can then format or arrange this text into another structure.&#x20;

Text processing is not an easy task to handle if we rely solely on programming languages.&#x20;

The emergence of LLMs can help us significantly with tasks involving text processing, including formatting.

### How it work ?

By formatting the text, we can clarify the structure of the output as we desire.&#x20;

If we do not specify the structure, the LLM will determine the most general output structure for all audiences, which may not be precise for domain-specific needs.&#x20;

The best approach is to provide detailed information about the desired structure of the output.

{% code overflow="wrap" %}

```
"In a quiet village, a young girl named Mia discovered a hidden key in her grandmother's attic. Curious, she followed a map etched on the key, leading her to an ancient oak tree in the forest. As she turned the key in a concealed lock, a door opened, revealing a magical world filled with talking animals and shimmering rivers. Mia befriended a wise fox who guided her through enchanting adventures. When she returned home, she knew the magic was real, for the key glowed warmly in her hand, a reminder of the wonders just beyond the ordinary."

Format the text into 3 sectors.
```

{% endcode %}

{% hint style="info" %}
A software developer could consider this ability to be a **new type** of capability for programming languages.
{% endhint %}

### Prompt example

{% embed url="<https://prompt.float16.cloud/prompt/347f008b-e02e-45f9-997b-324edcbb058b>" %}
CSV to JSON
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/8ed38a37-29fd-4939-bf70-18e777bbbe07>" %}
JSON to case report
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/e5dbe06a-95dd-4210-9097-5c5f574aecbe>" %}
User requirement to functional and non-functional requirement
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/99b65ed2-7e87-4142-9160-77f791ffcac5>" %}
Formatting text into 3 sectors
{% endembed %}


# Chat

### Intro

ChatGPT is one of the well-known use cases that draws our attention to the capabilities of LLMs.

### How it work ?

The chat capability of an LLM is not a magic trick. It's based on a prompt template that segments user turns and assistant (LLM) turns and combines them. That's why an LLM can recall information from previous chats.

The system prompt can be an important part because it allows us to control the chat style and chat objectives.

{% code overflow="wrap" %}

```
SYSTEM PROMPT
You are an assistant and need to collect data from the user, including their name, age, and gender. You need to guide and help the user input the correct data.
USER PROMPT
What about pricing ?
```

{% endcode %}

{% hint style="info" %}
A software developer could consider this ability to be a **new type** of capability for programming languages.
{% endhint %}

### Prompt example

{% embed url="<https://prompt.float16.cloud/prompt/b3f3884a-ff37-42aa-bc84-de263bcfd108>" %}
Change previous number into another language
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/9a265806-eeb7-4dd7-9c8c-cdceaee969fd>" %}
Chatbot to collect information
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/e4925c87-2261-4877-9168-9871a4459104>" %}
Chatbot to collect information and refused non-relevant task
{% endembed %}


# Technical term (Retrieve)

### Intro

An LLM can follow instructions via prompts. However, an LLM is more sensitive to "technical terms" rather than "general terms". The best way to get precise output is to use technical terms instead of general terms.

### How it work ?

We can change "general terms" into "technical terms" without adjusting other instructions.

````
``` python
def calculate_square_area (x, y) : 
  return x * y
```
Write unit test.
Return only code.
---------------------------------------------------------
``` python
def calculate_square_area (x, y) : 
  return x * y
```
Write boundary test.
Return only code.
````

{% hint style="info" %}
A software developer could consider this ability similar to the **search** functionality in other programming languages.
{% endhint %}

### Prompt example

{% embed url="<https://prompt.float16.cloud/prompt/42addd05-fdf9-40ee-9dd3-b7252349a806>" %}
General term to generate unit test
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/3074c4d8-7615-4e6d-b803-09a156033a72>" %}
Specific term to generate unit test
{% endembed %}

{% embed url="<https://prompt.float16.cloud/prompt/55becc6c-d3fa-4a7d-9ddc-6faa43ecba57>" %}
Specific term to generate unit test
{% endembed %}


# Privacy Policy

Float16's Privacy Policy

## Privacy Statement

**Effective Date:** 5th November 2024

**Table of Contents**

* [Introduction](#introduction)
* [Data Protection Officer](#data-protection-officer)
* [How we collect and use (process) your personal information](#how-we-collect-and-use-process-your-personal-information)
* [Legal Basis for Processing](#legal-basis-for-processing)
* [Purpose of Processing](#purpose-of-processing)
* [Data Retention Period](#data-retention-period)
* [Your Rights as a Data Subject](#your-rights-as-a-data-subject)
* [International Data Transfers](#international-data-transfers)
* [Security Measures](#security-measures)
* [Third-Party Processors](#third-party-processors)
* [Children's Privacy](#childrens-privacy)
* [Changes to Privacy Policy](#changes-to-privacy-policy)
* [Questions, concerns, or complaints  ](#questions-concerns-or-complaints)

### Introduction

The Float16 Co.,Ltd. is a Float16.cloud ("we," "us," or "our") operating a GPU-managed service platform that provides enterprise-grade GPU infrastructure management solutions. Our platform enables organizations to access and utilize GPU computing resources for artificial intelligence (AI) development and Large Language Model (LLM) applications. We handle various aspects of GPU infrastructure including deployment, scaling, maintenance, resource quotas, and resource allocation on behalf of our users.

As a technology service provider, we understand that you are aware of and care about your own personal privacy interests, and we take that seriously. This Privacy Policy explains how we collect, use, disclose, and protect your personal data in accordance with Thailand's Personal Data Protection Act B.E. 2562 (PDPA) and other applicable laws.

### Data Protection Officer

Float16 Co.,Ltd. is headquartered in 168/67 Village No. 5, Sai Noi Subdistrict, Sai Noi District, Nonthaburi Province 11150, in Thailand. Float16 Co.,Ltd. has appointed an internal data protection officer for you to contact if you have any questions or concerns about Float16 Co.,Ltd.’s personal data policies or practices. If you would like to exercise your privacy rights, please direct your query to Float16 Co.,Ltd.’s data protection officer. Float16 Co.,Ltd.’s data protection officer’s name and contact information are as follows:

> **Weerasak  Suwannapong**
>
> Float16 Co.,Ltd.
>
> 168/67 Village No. 5, Sai Noi Subdistrict, Sai Noi District, Nonthaburi Province 11150
>
> <weerasak.suw@float16.cloud>
>
> +66 97 130 5131

### How we collect and use (process) your personal information

Float16 Co.,Ltd. collects personal information about its website visitors and customers. With a few exceptions, this may include but not limited to:

**Identity and Contact Information:**

* Full name
* Work email address
* Work phone number
* Employer name
* Work address

**Technical Data:**

* IP address
* Browser type and version
* Operating system
* Website usage data
* Service usage statistics

**Service Data:**

* Program usage history
* Service-related communications
* Customer support interactions
* Service preferences

We use this information to provide prospects and customers with services.&#x20;

We do not sell personal information to anyone and only share it with third parties who are facilitating the delivery of our services.

From time to time, Float16 Co.,Ltd. receives personal information about individuals from third parties. Typically, information collected from third parties will include further details on your employer or industry. We may also collect your personal data from a third party website (e.g. Google workspace)&#x20;

### Legal Basis for Processing

We process your personal data under the following legal bases as per PDPA requirements:

* Consent: Where you have given explicit consent for specific purposes
* Contractual Necessity: To fulfill our contractual obligations
* Legal Obligation: To comply with legal requirements
* Legitimate Interests: Where processing is in our legitimate business interests and does not override your rights

### Purpose of Processing

We collect and use your personal data for the following purposes:

1. Service Delivery:

* Providing GPU infrastructure management services
* Managing user accounts and access
* Technical support and customer service

2. Service Improvement:

* Analyzing usage patterns
* Improving platform performance
* Developing new features

3. Communication:

* Service updates and notifications
* Technical announcements
* Customer support responses

4. Legal Compliance:

* Meeting regulatory requirements
* Responding to legal requests
* Maintaining security records

### Data Retention Period

We retain your personal data for:

* Active customers: Duration of the service relationship plus 3 years
* Prospective customers: 2 years from last interaction
* Technical logs: 1 year from creation
* Legal requirements: As required by applicable laws

### Your Rights as a Data Subject

Under the PDPA, you have the following rights:

* Right to be informed
* Right to access your personal data
* Right to data portability
* Right to object to processing
* Right to erasure ("right to be forgotten")
* Right to restriction of processing
* Right to rectification
* Right to withdraw consent

To exercise these rights, please contact our Data Protection Officer.

### International Data Transfers

We store and process data in Thailand and the United States. When transferring data internationally, we ensure:

* Adequate data protection measures
* Appropriate safeguards through standard contractual clauses
* Compliance with PDPA requirements for international transfers

### Security Measures

We implement appropriate technical and organizational measures including:

* Encryption of data in transit and at rest
* Access controls and authentication
* Regular security assessments
* Staff training on data protection
* Incident response procedures

### Third-Party Processors

We use the following main third-party processors:

* AWS (Cloud infrastructure)
* Supabase (Database management)
* Google Workspace (Business operations)
* Google Analytics (Website analytics)

All processors are bound by data processing agreements compliant with PDPA requirements.

### Children's Privacy

We do not knowingly collect or process personal data from children under 20 years of age without parental consent as required by the PDPA.<br>

### Changes to Privacy Policy

We reserve the right to update this Privacy Policy. Any changes will be posted on our website with an updated effective date.

### Questions, concerns or complaints

If you have questions, concerns, complaints, or would like to exercise your rights, please contact us at:

> **Float16 Co.,Ltd.**
>
> 168/67 Village No. 5, Sai Noi Subdistrict, Sai Noi District, Nonthaburi Province 11150
>
> <support@float16.cloud>
>
> <https://float16.cloud/>
>
> +66 804934501


# Terms & Conditions

Float16's Terms and Conditions

## TERMS OF SERVICE

Last Updated: 5th November 2024

PLEASE READ THESE TERMS OF SERVICE CAREFULLY. THIS IS A LEGALLY BINDING AGREEMENT.

### AGREEMENT TO TERMS

These Terms of Service constitute a legally binding agreement made between you, whether personally or on behalf of an entity ("you") and Float16 Co., Ltd. ("we", "us", or "our"), concerning your access to and use of our website and services, including any media form, media channel, mobile website or mobile application related, linked, or otherwise connected thereto (collectively, the "Site"). We are registered in Thailand and have our registered office at 168/67 Village No. 5, Sai Noi Subdistrict, Sai Noi District, Nonthaburi Province 11150.

### PAYMENT AND SERVICE TERMS

#### **Payment Models**

We offer two payment models for our services:

1. **Top-up System**

* Users can prepay credits into their account
* Credits are deducted based on actual usage
* Remaining credits have no expiration date
* Usage history and balance are accessible through the dashboard
* Minimum top-up amount: 10 USD
* Payment processing through Stripe

2. **Subscription**

* Available through direct contact: <support@float16.cloud> or +66 876604840
* A pricing based on a requirements
* It is an annual/semi-annual advance payment depending on the agreement.
* Includes dedicated support (on work day)
* Enterprise features available
* Customizable service level agreements (SLAs)

#### Refunds and Cancellation

**Top-up System**

* Unused credits are refundable with a 3% processing fee
* Refund requests must be submitted in writing
* Processing time: 15-30 business days
* Minimum refund amount: 10 USD

**Subscription**

* Cancellation period within 7 days from the date of service start, you can get a full refund if the service is not used.
* 30-day advance notice required for cancellation
* A pro-rata refund may be requested for unused full periods at least 1 month, less a handling fee of up to 30%, depending on the agreement.
* No refund for partial months

### LEGAL COMPLIANCE

#### Thai Law Compliance

These Terms comply with:

* The Civil and Commercial Code of Thailand
* The Electronic Transactions Act B.E. 2544 (2001)
* The Consumer Protection Act B.E. 2522 (1979)
* The Personal Data Protection Act B.E. 2562 (2019)

#### Tax Considerations

* All prices are exclusive of VAT
* VAT will be charged where applicable under Thai law
* Tax invoices will be issued electronically
* Users are responsible for any withholding tax obligations

### SERVICE USAGE AND LIMITATIONS

#### Fair Usage Policy

* API rate limits apply
* Concurrent connection limits may apply
* Resource usage monitoring and throttling may be implemented
* Automated access must be pre-approved

#### Service Level Agreement

**Top-up System**

* Best-effort support
* Standard API availability: 99%
* Response time: within 48 office hours

**Subscription**

* Enhanced SLA available
* Custom uptime guarantees
* Priority support
* Dedicated account manager

### DATA PROTECTION AND PRIVACY

In compliance with Thailand's Personal Data Protection Act (PDPA):

#### **Data Collection and Processing**

* We collect and process data as specified in our Privacy Policy
* Data is processed in accordance with Thai law
* Users have rights to access, correct, and delete their data

#### Data Security

* Industry-standard encryption
* Regular security audits
* Data breach notification within 72 hours
* Cross-border data transfers comply with PDPA requirements

### INTELLECTUAL PROPERTY

#### Ownership

* All content and materials remain our property
* Usage rights are non-transferable
* Licensed on a non-exclusive basis

#### User Content

* You retain ownership of your content
* You grant us license to process and store your content
* Content must not infringe third-party rights

### TERMINATION

We reserve the right to terminate services:

* For violation of these Terms
* For illegal or unauthorized use
* For non-payment
* For extended periods of inactivity

### DISPUTE RESOLUTION

#### Informal Resolution

* Parties will attempt to resolve disputes informally
* 30-day negotiation period

#### Formal Resolution

* Disputes shall be resolved in Thai courts
* Jurisdiction: Courts of Thailand
* Thai law governs these Terms

### LIMITATIONS OF LIABILITY

To the extent permitted by Thai law:

* We limit liability to direct damages
* Maximum liability limited to fees paid
* No liability for:
  * Service interruptions
  * Data loss
  * Consequential damages
  * Third-party actions

### CONTACT INFORMATION

Float16 Co., Ltd.

* Address: 168/67 Village No. 5, Sai Noi Subdistrict, Sai Noi District, Nonthaburi Province 11150
* Email: <support@float16.cloud>
* Phone: +66 804934501

### AMENDMENTS

We reserve the right to modify these Terms:

* With 30 days notice for material changes
* Changes effective immediately for non-material updates
* Notice provided via email or Site announcement
* Continued use constitutes acceptance

***

By using our services, you acknowledge that you have read, understood, and agree to be bound by these Terms of Service.

<br>


