# Cloud Provider IAM Authentication Source: https://docs.freeplay.ai/account-setup/cloud-provider-iam-auth Configure IAM-based authentication for Amazon Bedrock and Google Vertex AI models. By default, Freeplay routes requests to LLM providers using API keys or credentials managed by Freeplay. For customers with compliance, security, or access requirements, Freeplay also supports IAM-based authentication flows that let you use your own cloud provider roles and service accounts. This page covers IAM authentication setup for **Amazon Bedrock** (via AWS assume role) and **Google Vertex AI** (via GCP service account impersonation). **BYOC** customers skip to the BYOC section below. *** ## Amazon Bedrock \*\*Note: \*\*BYOC customers, see next section. When no custom credentials are configured, Freeplay uses its own AWS credentials to call Bedrock models in the Freeplay VPC. ### Using your own AWS role If you need Freeplay to call models in your own AWS account — for compliance reasons, to access private models, or to maintain full control over credentials — you can configure [AWS assume role authentication](https://docs.aws.amazon.com/STS/latest/APIReference/API_AssumeRole.html). With assume role auth, Freeplay temporarily assumes an IAM role in your AWS account using a short-lived token. You retain full control and can revoke access at any time. #### How it works 1. Freeplay authenticates with an internal AWS role 2. That role assumes **your** AWS role, validated by a shared External ID 3. Using the resulting short-lived token, Freeplay calls Bedrock models in your account The External ID is a shared string you configure both in your AWS trust policy and on the Freeplay **Settings > Models** page. It prevents unauthorized parties from assuming your role. #### Customer setup Freeplay strongly recommends creating a dedicated, isolated IAM role specifically for Freeplay to access Bedrock. **Step 1: Create an IAM role for Freeplay** Create a new IAM role in your AWS account with a trust policy that allows Freeplay's role to assume your role, validated by an External ID. Contact your Freeplay account team to obtain the Freeplay role ARN to use as the **Principal** in your trust policy. You will also need to generate a unique External ID string — this same string must be configured both in your trust policy and on the Freeplay **Settings > Models** page. **Step 2: Attach a permissions policy** Attach the following permissions policy to the role to grant Freeplay access to invoke Bedrock models: ```json theme={null} { "Version": "2012-10-17", "Statement": [ { "Sid": "AllowBedrockInvoke", "Effect": "Allow", "Action": [ "bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream" ], "Resource": "*" } ] } ``` You can scope the `Resource` field to specific model ARNs if you want to restrict which Bedrock models Freeplay can invoke. **Step 3: Configure in Freeplay** 1. Navigate to **Settings > Models** in Freeplay 2. Under Amazon Bedrock, enter: * Your IAM role ARN (e.g., `arn:aws:iam:::role/`) * The External ID you set in the trust policy 3. Mark this authentication method as default for the provider 4. Save your configuration Freeplay will now use assume role authentication for all Bedrock requests. ### BYOC (Bring Your Own Cloud) setup For customers running Freeplay in their own VPC via [BYOC deployment](/security-compliance/byoc), the authentication flow is simplified: 1. Freeplay uses an implicit role from [AWS IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html) (IAM Roles for Service Accounts) attached to the pod — no Freeplay-managed user or role is involved 2. The IRSA role assumes your configured AWS role to retrieve a short-lived token Customer setup follows the same steps as above (create a role, attach the permissions policy, configure in Freeplay), with one difference: * The **Principal** in your trust policy should reference the IRSA role ARN provided during your BYOC onboarding, rather than the standard Freeplay role ARN * The External ID condition is optional but recommended for BYOC deployments *** ## Google Vertex AI ### Default behavior When no custom credentials are configured, Freeplay uses its own GCP service account to call Vertex AI models. ### Using your own GCP project To route Vertex AI requests through your own GCP project, you need to grant Freeplay's service account permission to create access tokens in your project. #### Customer setup **Step 1: Grant the Service Account Token Creator role** 1. In the GCP project where your Vertex AI models are hosted, go to **IAM & Admin > IAM** 2. Click **Grant Access** at the top of the page 3. Set the following: * **Principal**: Contact your Freeplay account team for the service account email to use * **Role**: Service Account Token Creator 4. Click **Save** Role changes can take 30 seconds to 2 minutes to propagate in GCP. If you see permission errors immediately after saving, wait a moment and try again. **Step 2: Configure in Freeplay** 1. Navigate to **Settings > Models** in Freeplay 2. Configure your Vertex AI provider settings with your GCP project details 3. Mark this authentication method as default for the provider 4. Save your configuration Freeplay will now use service account impersonation to call Vertex AI models in your project. # Configure LiteLLM Proxy Models in Freeplay Source: https://docs.freeplay.ai/account-setup/configure-litellm-proxy-models-in-freeplay Set up LiteLLM Proxy as a custom provider to access multiple models through a unified interface. * **Model Flexibility**: Easily switch between different LLM providers while using only OpenAI code for all model interactions. * **Unified Interface**: Use a consistent API format regardless of the underlying model * **Simplified Management**: Access numerous models through a single integration * **Custom Models**: Reference custom-deployed LLMs with the same workflow ## Setting Up LiteLLM Proxy in Freeplay ### Step 1: Set up LiteLLM Proxy & Add Models In this example we will use gpt-3.5-turbo and claude-3-5-sonnet via Anthropic. To get set up with LiteLLM you can see the full docs [here](https://docs.litellm.ai/docs/proxy/docker_quick_start). To start, ensure you have a model config file like the one below: ```yaml yaml theme={null} model_list: - model_name: gpt-3.5-turbo # Use this exact name in Freeplay litellm_params: model: openai/gpt-3.5-turbo api_key: os.environ/OPENAI_API_KEY - model_name: claude-3-5-sonnet # Use this exact name in Freeplay litellm_params: model: anthropic/claude-3-5-sonnet-20240219 api_key: os.environ/ANTH_API_KEY ``` **Important**: When configuring models, the name in Freeplay must match the `model_name` in your LiteLLM Proxy configuration file. ### Step 2: Configure LiteLLM Proxy as a Custom Provider 1. Navigate to Settings in your Freeplay account 2. Find "Custom Providers" section 3. Enable "LiteLLM Proxy" provider image.png ### Step 3: Add an API Key 1. Click "Add API Key" 2. Name your API key 3. Enter your LiteLLM Proxy Master key 4. Optionally, mark it as the default key for LiteLLM Proxy image.png ### Step 4: Add Your LiteLLM Proxy Models 1. Select "Add a New Model" 2. Optionally select if the model supports tool use 3. Optionally add a display name for the model 4. Enter link to your LiteLLM API Proxy Note: Token pricing information is automatically fetched from LiteLLM Proxy so you do not need to provide it. image.png ### Step 5: Using LiteLLM Proxy Models in the Prompt Editor 1. Open the prompt editor in Freeplay 2. In the model selection field, search for "LiteLLM Proxy" 3. Select one of your configured models image.png ## Integrating LiteLLM Proxy with Your Code The following example shows how to configure and use LiteLLM Proxy with Freeplay in your application. The benefit of using LiteLLM Proxy is you only need to configure your calls to work with OpenAI, LiteLLM Proxy will handle all the formatting: ```python python theme={null} ####################### ## Configure Clients ## ####################### # Configure the Freeplay Client fp_client = Freeplay( freeplay_api_key=API_KEY, api_base=f"{API_URL}" ) # Call OpenAI using the LiteLLM url and api_key. # This handles the routing to your models while keeping the response # in a standard format. client = OpenAI( api_key=userdata.get("LITE_LLM_MASTER_KEY"), base_url=userdata.get("LITE_LLM_BASE_URL") ) ##################### ## Call and Record ##################### # Get the prompt from Freeplay formatted_prompt = fp_client.prompts.get_formatted( project_id=PROJECT_ID, template_name=prompt_name, environment=env, variables=prompt_vars, history=history ) # Call the LLM with the fetched prompt and details start = time.time() completion = client.chat.completions.create( messages=formatted_prompt.llm_prompt, model=formatted_prompt.prompt_info.model, tools=formatted_prompt.tool_schema, **formatted_prompt.prompt_info.model_parameters ) # Extract data from LiteLLM response completion_message = completion.choices[0].message tool_calls = completion_message.tool_calls text_content = completion_message.content finish_reason = completion.choices[0].finish_reason end = time.time() print("LLM response: ", completion) # Record to Freeplay ## First, store the message data in a Freeplay format updated_messages = formatted_prompt.all_messages(completion_message) ## Now, record the data directly to Freeplay completion_log = fp_client.recordings.create( RecordPayload( project_id=PROJECT_ID, all_messages=updated_messages, inputs=prompt_vars, session_version_info=session, trace_info=trace, prompt_info=formatted_prompt.prompt_info, # Note: you must pass UsageTokens for the cost calculation to function call_info= CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end, UsageTokens(completion.usage.prompt_tokens, completion.usage.completion_tokens)) ) ) ``` ## Current Limitations * **Automatic Cost Calculation:** UsageTokens must be passed for cost calculations to work with LiteLLM Proxy. Also, for self hosted models that depend on time, cost calculation is not currently supported. * **Auto-Evaluation Compatibility**: Model-graded evals that are configured and run by Freeplay do not currently support LiteLLM Proxy models. ## Additional Resources * [LiteLLM Proxy Server Documentation](https://docs.litellm.ai/docs/proxy/quick_start) * [LiteLLM Supported Models](https://docs.litellm.ai/docs/providers) * [Configure, Test & Deploy a Fallback LLM Provider](/practical-guides/configuring-a-fallback-llm-provider-with-freeplay) * [Voice-Enabled AI with Pipecat, Twilio, and Freeplay](/practical-guides/build-voice-enabled-ai-applications-with-pipecat-twilio-and-freeplay) ``` ``` # Model Management Source: https://docs.freeplay.ai/account-setup/model-management Configure LLM providers, manage API keys, and control which models your team can use. Only Freeplay admin role users have permission to manage model access and keys. *** ## Configuring Model Access Freeplay has first-class support for calling models from common hosts/providers including OpenAI, Anthropic, Azure OpenAI Service, Amazon Bedrock, Amazon SageMaker, Groq, Baseten and more. These providers can be directly configured in the Freeplay UI, including configuring appropriate endpoints and API keys or other relevant credentials. These models can then be used end-to-end in the Freeplay application, including in our playground UI. At the same time, **our SDKs allow you to call any model you want** and record the results with Freeplay. The Freeplay application then lets you configure those models as part of your prompt templates and experiments. An SDK example of calling other models is [here](/freeplay-sdk/recording-completions#calling-any-model). You can control which of these models and providers your team is able to use and deploy on the Models page. Navigate to **Settings > Models** to configure models for your team. * You can disable Default models (e.g. if your team doesn't have permission to use a given provider) * You can add your own models and endpoints to default providers, like OpenAI fine-tuned models or Llama 3 on SageMaker * You can control configurability on prompt templates for any other custom models you might have logged with Freeplay *** ## API Keys & Credentials ### Bringing Your Own Keys to Freeplay Freeplay allows customers to store their LLM provider keys with Freeplay so that Freeplay's application can route requests in the interactive Prompt Editor and for UI-driven Tests. LLM provider keys stored with Freeplay are only used for routing LLM requests, and will not be surfaced in the Freeplay dashboard or via the Freeplay API. ### Security Matters Freeplay uses application-level encryption to encrypt customer LLM provider keys both at rest and in transit. Keys are only decrypted prior to routing requests to customer models. This level of encryption is supplementary to transparent data encryption provided by cloud providers. We follow industry best practices of encryption key management including regular key rotation and audit logging of all key access. More details on security [here](/security-compliance/security-overview). ### Customer Key Best Practices 1. Provide a unique API key with finely-scoped access for use by Freeplay. Freeplay's application only needs access to make requests to the inference endpoints for your provider. * e.g. for OpenAI, we recommend creating a separate Project with a cost limit, and creating a key with only access to the `/v1/chat/completions` endpoint. 2. Use Freeplay's UI to rotate your LLM provider keys following your organization's guidelines. As soon as a key is updated in Freeplay's UI, the old value will be destroyed, and the new value will be used for future requests. 3. Monitor usage of your API keys regularly. Freeplay does not set limits on use of customer keys other than those imposed by the LLM providers. ### Configuring Keys in the Dashboard You can set a default API key for a given provider by clicking on the provider name in the list. This default key will be used for most situations. For advanced use, you can also link different keys to each endpoint for a given provider, e.g. if you want to use a different key for a fine-tuned model. Set a different key by clicking on the model name. ### Isolating API keys for test runs Freeplay [test runs](/core-concepts/test-runs/test-runs) execute evaluations with high parallelism to deliver fast results. Without proper key isolation, this parallel traffic can consume rate limits shared with your production application. To prevent this, create a dedicated API key for Freeplay that is scoped to its own rate limits, separate from your production keys. This ensures test run traffic never competes with your live application for capacity. If you use the same API key for both Freeplay and your production application, parallel test runs may trigger rate limits (HTTP 429 errors) that affect your production traffic. Freeplay handles 429 responses with exponential backoff, but this does not protect your production application from hitting the same shared limits. OpenAI supports project-level API keys with independent rate limits: 1. Go to [platform.openai.com](https://platform.openai.com) → **Settings** → **Projects** 2. Click **Create project** (e.g., "Freeplay") 3. Under the new project, go to **API Keys** → **Create a new key** 4. Go to the project's **Limits** tab and set TPM/RPM caps you are comfortable allocating to Freeplay (e.g., 50% of your org limit) 5. In Freeplay, navigate to **Settings** → **Models** and update your OpenAI key to the new project-scoped key Anthropic supports workspace-level API keys with independent rate limits: 1. Go to [console.anthropic.com](https://console.anthropic.com) → **Settings** → **Workspaces** 2. Create a new workspace (e.g., "Freeplay") 3. In the workspace, go to **API Keys** → **Create a new key** 4. Go to the workspace's **Limits** tab and set caps 5. In Freeplay, navigate to **Settings** → **Models** and update your Anthropic key to the workspace-scoped key You cannot set custom rate limits on Anthropic's default workspace. You must create a new workspace to configure independent limits. If your provider is not listed above, the same principle applies: create a separate API key dedicated to Freeplay with its own rate limits or usage scope. This prevents test run traffic from affecting your production application. Check your provider's documentation for options like projects, workspaces, or service accounts that support independent rate limiting. *** ## Provider-Specific Configuration For IAM-based authentication with **Amazon Bedrock** or **Google Vertex AI** (using your own AWS roles or GCP service accounts instead of API keys), see [Cloud Provider IAM Authentication](/account-setup/cloud-provider-iam-auth). ## Configuring OpenAI Fine-Tuned Models You can configure fine-tuned OpenAI models by navigating to **Settings > Models** and choosing **Add fine-tuned model** under the OpenAI Details heading. Enter the name of your fine-tuned model as provided by OpenAI. You may also optionally enter a more readable display name and specify an associated API key. ### Calling Your Fine-Tuned Model When creating or editing Prompt Templates you will now see **fine-tuned** as an option in the Model dropdown. Model Version will then populate with all your fine tuned models. You can call your fine-tuned model from within the prompt editor as well as use your fine-tuned model in the [Freeplay SDK](/freeplay-sdk/recording-completions#record-an-llm-interaction) just as you would any other OpenAI model. Keying the model name, messages and model parameters off of the prompt object. # Project Setup Source: https://docs.freeplay.ai/account-setup/project Create your Freeplay account, generate API keys, and configure models and environments. ## Create your account Start by signing up for Freeplay at app.freeplay.ai. Once you've created your account, you'll land in your workspace where you can create projects, configure models, and manage your team. ### Custom Freeplay Subdomains If you have a custom subdomain, it will act as a unique identifier that links your SDK to your Freeplay instance. Here's how to find and use it: * Your subdomain is part of your Freeplay URL. For example, if your Freeplay URL is `https://acmecorp.freeplay.ai`, then your subdomain is `acmecorp`. * Ensure you input this subdomain correctly in your SDK configuration. It's crucial for directing your SDK's requests to the right instance. *** ## Generate Freeplay API Keys Freeplay supports user scoped and project scoped API keys. User based keys are created at the account settings level and inherit all user permissions. Project scoped keys represent service accounts and are scoped to that project specifically. ### To create a User API key: 1. Navigate to Settings > API Access in your Freeplay dashboard 2. Click "Create API Key" 3. Give your key a descriptive name (e.g., "Production" or "Development") 4. Click the copy button to copy your full API key 5. Store it securely—you won't be able to see it again ### To Create a Project API Key First, you must have a Freeplay project and navigate to the projects settings. Then as an admin you can: 1. Select "Service Accounts" 2. Create a new service account, provide it a name 3. Create an api key for this service account For more details on API keys and user roles, see our RBAC guide [here](/account-setup/role-based-access-control). *** ## Configure Models Before you can start building prompts, you'll need to configure which AI models your team can use. Freeplay might include starter credits to help you explore, but we recommend adding your own API keys from providers like OpenAI or Anthropic for ongoing use. #### Basic setup: Navigate to Settings > Models to see available providers and models. By default, you'll see common providers like OpenAI, Anthropic, and others. If you've already added API keys for these providers, you're ready to start building prompts. #### What you can configure: * Enable or disable specific models and providers * Add your LLM provider API keys for use in the playground and tests * Configure custom endpoints (e.g., Azure OpenAI, fine-tuned models) * Control which models your team can deploy to production Need more control? Check out our detailed [Model and Key Management guide](/account-setup/model-management) for more information. *** ## Set Up Environments Freeplay lets you deploy different prompt versions across multiple environments, making it easy to follow a traditional promotion flow from development to production. Default environments: * latest - Automatically assigned to new prompt versions * production - Your stable, live version * sandbox - For testing before production * dev - Development environment *** ## Freeplay Project ID Your Project ID is what connects SDK logs to the right project in Freeplay. To find it, just navigate to your project and look at the URL—the long string of characters after /projects/ is your Project ID. For example: `https://app.freeplay.ai/projects//`. We recommend that you store this ID as an env variable and reference it as `FREEPLAY_PROJECT_ID`. *** ## Project access settings Projects can be configured as **public** (accessible to all users in your organization) or **private** (accessible only to specific invited members). Private projects are useful for sensitive data that only a subset of team members should access. To change project access settings, go to **Edit Project > Project access**. For more details on project visibility and user permissions, see [Private vs. Public Projects](/account-setup/role-based-access-control#private-vs-public-projects). *** With your account configured, API key ready, and models set up, you're ready to start building with Freeplay. Head to the [Quick Start guide](/getting-started/overview) to create your first prompt and run your first test. # User Roles and Access Controls Source: https://docs.freeplay.ai/account-setup/role-based-access-control Manage team permissions with role-based access controls for users and projects. ## Roles There are 4 primary roles in the Freeplay product. ### Admins Admins in Freeplay can perform any action, including inviting and managing other users and configuring models available to your team. Use the admin role for those who will be spearheading the adoption of Freeplay within your organization. We recommend having at least 2-3 admins on your account, large organizations may benefit from having additional admins. ### Users Users can perform most actions in Freeplay but are primarily restricted from taking actions which could have major impacts on the overall account, particularly as it relates to compliance and user management. Use this role for those who are contributing heavily to day to day development. Users are *restricted* from the following actions: * Inviting other users * Managing other users * Configuring account-level LLM providers and models * Setting account-level spend limits ### Analysts Analysts inherit the restrictions of Users but are additionally prohibited from performing sensitive actions like prompt & model deployment or data deletion. Use this role for those who will be reviewing data, running experiments, and evaluating data, but who won't be actively pushing changes to a production system. In addition to the aforementioned restrictions of Users, Analysts are also *restricted* from the following actions: * Deploying prompt templates * Deleting prompt templates * Deleting production data * Deleting dataset examples * Managing API keys (create or delete) * Managing environments (create or delete) * Configuring models (at the account level) * Project creation ### Guests Guests have the same permissions as Analysts, except that they can only see data for the specific projects they are invited to. (Unlike Analysts who have account-level/site-wide permissions to access any shared projects.) This role is intended for contractors or other collaborators with a narrow scope of focus. ## Role Enforcement & Project-Level Permissions Each user will be given a role at the account level. This will determine what the user can and can't do by default in any project (i.e. site-wide, default permissions). However, a user's role can be changed within the context of a specific project. For example, an account level Analyst can be given an Editor role on a specific project. The user will have all the permission associated with an Editor in that specific project, but will maintain their Analyst role at the account level and across other projects. ## Private vs. Public Projects There are two levels of accessibility for projects. Project privacy is determined at creation, but can also be changed after creation. To change this setting on a project, go to **Edit Project > Project access** and select either **"Anyone at your organization"** or **"🔒 Only specific members"**. **Public Projects** Public projects can be accessed by all users on the account with their default user roles. Users do not need to be specifically granted access to the project in order to access it (except Guests, who must be added). We recommend using this project type for any projects that do not contain sensitive data. **Private Projects** Private projects can only be accessed by users who have been directly granted access to the project. Each private projects will have 1 or more project admin who determine access for the project. We recommend using this project type for any projects that contain data that only a subset of internal employees or contractors are authorized to see. ## Service Accounts **Note:** Only account administrators can create, modify, or delete service accounts and their associated API keys. Service accounts enable project-scoped API key management for production environments. Each project can have multiple service accounts, and each service account can have multiple API keys, allowing you to separate keys across different environments (such as production and testing) or data sources. To configure service accounts, navigate to your project and select Service Accounts. From there, you can create new service accounts and generate API keys for each one. # Single Sign-On and SCIM Source: https://docs.freeplay.ai/account-setup/sso-and-scim Enterprise authentication options for centralized user management and automated provisioning. Freeplay offers enterprise-grade authentication options to help organizations manage access at scale. This page covers Single Sign-On (SSO) via SAML and automated user provisioning through SCIM. SSO and SCIM are available for **Enterprise** tier customers only. Contact [support@freeplay.ai](mailto:support@freeplay.ai) to enable these features for your account. ## Default Authentication Options By default, Freeplay supports two authentication methods: * **Email and password** — Standard username/password authentication * **Google Workspace SSO** — Sign in with your Google account ## Single Sign-On (SSO) via SAML Enterprise customers can enable SAML-based Single Sign-On to authenticate users through their organization's Identity Provider (IdP). This allows your team to use their existing corporate credentials to access Freeplay. ### Supported Identity Providers Thanks to our authentication partner WorkOS, Freeplay supports most major Identity Providers including: * Okta * Microsoft Entra ID (Azure AD) * Cisco Duo * OneLogin * JumpCloud * And many others ### Enabling SSO To enable SSO for your organization: 1. **Contact Freeplay** — Reach out to [support@freeplay.ai](mailto:support@freeplay.ai) to request SSO enablement 2. **Self-serve configuration** — We'll send your IT/auth administrator an email invite to configure SSO through WorkOS 3. **Complete setup** — Your admin completes the SAML configuration in your IdP 4. **Activation** — Once SSO is enabled, other authentication methods (email/password, Google) are disabled for your account 5. **Account creation & role management** — You will continue to add/remove users and update roles manually via the Freeplay UI unless you choose to additioanlly enable SCIM (see below) Once SSO is enabled, users will only be able to authenticate through your organization's Identity Provider. Make sure your IdP configuration is correct before completing the transition. ## Automated User Provisioning with SCIM Freeplay offers SCIM (System for Cross-domain Identity Management) support for Enterprise customers to automate the provisioning and deprovisioning of user accounts. This allows you to manage Freeplay user access directly from your Identity Provider (IdP). ### Overview Enabling SCIM/Directory Sync streamlines your user management by treating your directory as the single source of truth. * **Automated Onboarding** — New users assigned to Freeplay groups in your IdP are automatically created in Freeplay * **Automated Offboarding** — Deactivating a user in your IdP immediately revokes their access to Freeplay * **Centralized Role Management** — User roles are determined solely by their group membership in your directory ### Prerequisites and Setup To enable SCIM for your account: 1. **Contact Support** — Reach out to the Freeplay team to request SCIM enablement. Please provide the email address of the person who will handle the configuration. This should be somebody that is appropriately permissioned to manage your IdP. 2. **Configuration** — We will send an invite link to configure your directory sync via WorkOS (our third-party auth provider). This link will be in the form of `https://setup.freeplay.ai/init?setupLinkToken=`. 3. **Integration** — Once the configuration is done, either synchronously or async, let us know that you're ready to schedule a cutover time. We'll verify the details and prep your account for the switch for that time. 4. **Transition/Cutover** — Once SCIM is enabled for your account, it will no longer be possible to add or manage users in the UI. We recommend being ready to test quickly with a few sample users to verify that roles are mapping correctly. ### Role Mapping Freeplay uses a strict **1:1 mapping** between your directory groups and Freeplay user roles. To assign a role to a user, you must add them to **one and only one** of the specific groups listed below. If a user is in multiple Freeplay groups at a single time, it is undefined which role from that set they will assume. #### Default Supported Group Names If you want a more seamless setup for role management, you can name your IdP groups like the following and we will directly sync the users into Freeplay via a strict string match on these values. | IdP Group Name | Freeplay Role | | ------------------ | ------------- | | `freeplay_admin` | Admin | | `freeplay_user` | User | | `freeplay_analyst` | Analyst | | `freeplay_guest` | Guest | #### Custom Group Names If you prefer to define your own group names, the 1:1 mapping still applies. You will need to provide us with a list of your relevant IdP groups and how each one should map to Freeplay's user roles (Admin, User, Analyst, Guest). This is a manual process today. Changes to your group/group names without informing us may result in users being demoted to the Analyst (default role). Use [support@freeplay.ai](mailto:support@freeplay.ai) for any desired changes here. For example: | IdP Group Name | Freeplay Role | | ----------------- | ------------- | | `MyCoApp-admin` | Admin | | `MyCoApp-dev` | User | | `MyCoApp-analyst` | Analyst | | `MyCoApp-vendor` | Guest | Users that are synced to Freeplay with a group name that does not match one of the strings above (or with no group) will default to the **Analyst** role. ### Managing Access and Limitations Once Directory Sync is enabled for your account, user management behaviors change significantly. Please review the following rules to ensure a smooth workflow. #### Pre-existing Users * When first setting up the directory to sync with Freeplay, if you already have users in the Freeplay system they will be retained. As noted below, they will no longer be editable via the Freeplay app. * **Existing users will retain their assigned role.** In order to ensure your IdP is in sync with what's in Freeplay, you should manually add those existing users to the appropriate group in your IdP. * **To get a list of the users with their roles:** You can either get them from the UI or from the API. A call to `GET /api/v2//users` will return all active users in the account. Note that to call this endpoint the API key must be associated with a user with the Account Admin role. * Their role will change if they're added to other groups, or updated. #### Directory as Source of Truth * **No UI Creation** — You will no longer be able to create new users inside the Freeplay app UI. All users must be provisioned via SCIM. * **Locked Profile Data** — Users' First Name, Last Name, and Email are set during provisioning and cannot be edited within Freeplay. * **Permissions** — User roles are locked to their directory group. You cannot change a user's role manually in the Freeplay UI. #### Changing User Roles Freeplay determines a user's role based on **the most recent event received from your directory** (i.e. "last event wins"). We do not support mapping multiple groups to a single user. **To change a user's role (e.g., from User to Admin):** Role changes are no longer allowed inside the app. This is to prevent out-of-sync issues. Roles are now handled like so: 1. **Remove** the user from the old group (`freeplay_user`) **first** 2. **Add** the user to the new group (`freeplay_admin`) **second** If you add a user to a new group before removing them from the old one, the "remove" event for the old group may arrive last. This will leave the user with no valid group, causing them to revert to the default **Analyst** role. This can be easily adjusted by deactivating/reactivating that user's group setting. #### Project Membership SCIM manages **Roles** (system-level permissions), but it does not manage Freeplay **Project** membership. * **Private Projects** — Access to private projects must still be granted manually within the Freeplay UI by a Project admin. * **Guest Users** — Since the Guest role usually implies restricted access (e.g., for contractors), you must explicitly add them to specific Projects in the Freeplay UI for them to see any content. ### Technical Limitations * **Timing** — We process events on a 5-minute interval. This means that role changes can take up to 5 minutes to process. Please be aware that role changes will not be instantaneous. * **Single Group Only** — Users should belong to only one Freeplay-mapped group at a time. If a user is added to multiple groups, the role will reflect the most recent "Add" event received. This can also be fixed on your side by removing and re-adding the user to the highest privileged group that's desired. (e.g. if a user is added to IdP-User and IdP-Admin, but the user event arrived later, they'll be set as IdP-User. Simply re-add them to IdP-Admin to re-grant the admin permissions here). * **No `memberOf` Support** — We do not support the use of `memberOf` SAML assertions to update permissions on login. Permissions are updated only via SCIM directory events. # Token Cost Estimates Source: https://docs.freeplay.ai/account-setup/token-cost-estimates Understand how Freeplay calculates and displays token cost estimates to help you track and manage LLM API spend. ## Overview Cost estimates in Freeplay give you a clear view of how you are spending your API tokens. Freeplay breaks down costs by project, prompt template, and evaluation so you can trace your spend and identify where costs are highest. Freeplay also provides rough estimates of evaluation costs before you run them, helping you make informed decisions about your testing and evaluation strategy. ## Where to find cost estimates You can view cost information in two places within Freeplay. ### Account settings --> Usage The [**Usage** tab](https://app.freeplay.ai/settings/members) in Account Settings is the primary source of truth for costs and cost estimates. Navigate to **Account Settings** and select the **Usage** tab. This page provides two key metrics: * **App spend** - estimated spend based on logged token counts to Freeplay. This reflects the cost of running your application's LLM calls, not money spent in or through Freeplay itself. * **Spend via Freeplay** - the total spend from activity within the Freeplay platform. This includes: * Online evaluations and auto-categories * Test runs * Prompt playground usage * [AI features](/core-concepts/ai/ai-features) ### Estimated evaluation costs On the evaluations page, each evaluation displays an estimated cost. This estimate gives you a rough idea of how much it costs to run that evaluation over your completions. The estimate is calculated as follows: 1. Takes the average input and output tokens for the specific prompt template or agent (including input variables) from the last week 2. Calculates the average cost based on those token counts This is a rough estimate and is subject to change based on volume, sampling rate, and other factors. ## How costs are calculated Freeplay bases all cost estimates on token usage. Token costs are calculated using up-to-date pricing information from each provider. Cost estimates may vary if you: * Use different LLM providers * Have custom deployment configurations (e.g., Azure OpenAI, AWS Bedrock) * Have negotiated pricing agreements with providers ### Custom per-token costs For certain providers such as LiteLLM, you can provide custom input and output per-token cost information. This allows Freeplay to reflect your actual costs more accurately when you have special pricing or use self-hosted models. To configure custom token costs, see [LiteLLM Proxy](/account-setup/configure-litellm-proxy-models-in-freeplay). ## Frequently asked questions * **Model selection** -- switching to a smaller or less expensive model for tasks that do not require the most capable model * **Prompt optimization** -- reducing token counts by writing more concise prompts * **Sampling rate** -- adjusting the sampling rate for online evaluations to run them on a subset of completions rather than all of them * **Evaluation frequency** -- running test evaluations less frequently or on smaller datasets during development Freeplay uses publicly available pricing from each LLM provider to estimate the cost per input and output token. These rates are updated regularly to reflect the latest published pricing. If you use self-hosted models, custom deployments, or have negotiated pricing with providers, the default per-token costs may not reflect your actual spend. For providers that support it (such as LiteLLM), you can configure custom per-token costs in Freeplay to get more accurate estimates. For other providers, treat the estimates as a relative benchmark for comparing costs across prompts and models. # Bulk Create Agent Test Cases Source: https://docs.freeplay.ai/api-reference/agent-datasets/bulk-create-agent-test-cases https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/bulk Add multiple test cases to an agent dataset in a single request. Use for batch imports from production traces. Maximum 100 test cases per request. # Bulk Delete Agent Test Cases Source: https://docs.freeplay.ai/api-reference/agent-datasets/bulk-delete-agent-test-cases https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/bulk Remove multiple agent test cases in a single request. Use for batch cleanup operations. Maximum 100 test cases per request. # Create Agent-Level Dataset Source: https://docs.freeplay.ai/api-reference/agent-datasets/create-agent-level-dataset https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/agent-datasets Create a new dataset for agent-level testing. Use to organize test cases for evaluating full agent workflows. # Delete Agent Dataset Source: https://docs.freeplay.ai/api-reference/agent-datasets/delete-agent-dataset https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agent-datasets/{dataset_id} Archive an agent dataset and its test cases. Use when retiring datasets no longer needed. # Delete Agent Test Case Source: https://docs.freeplay.ai/api-reference/agent-datasets/delete-agent-test-case https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/{test_case_id} Remove a test case from an agent dataset. Use to clean up invalid or outdated test cases. # Get Agent Dataset Source: https://docs.freeplay.ai/api-reference/agent-datasets/get-agent-dataset https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets/{dataset_id} Retrieve an agent dataset's metadata by ID. Use to check dataset configuration or compatible agents. # Get Agent Test Case Source: https://docs.freeplay.ai/api-reference/agent-datasets/get-agent-test-case https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/{test_case_id} Retrieve a specific agent test case by ID. Use to inspect test case details or debug test run results. # List Agent Datasets Source: https://docs.freeplay.ai/api-reference/agent-datasets/list-agent-datasets https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets Retrieve all agent-level datasets in a project. Use to discover available datasets for agent test runs. `page_size` defaults to 30, maximum 100. # List Agent Test Cases Source: https://docs.freeplay.ai/api-reference/agent-datasets/list-agent-test-cases https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases Retrieve all test cases (row-level values) in an agent dataset. Use to review dataset contents or export for analysis. `page_size` defaults to 30, maximum 100. # Update Agent Dataset Source: https://docs.freeplay.ai/api-reference/agent-datasets/update-agent-dataset https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/agent-datasets/{dataset_id} Modify an agent dataset's name, description, or compatible agents. Use to evolve datasets as agent requirements change. # Update Agent Test Case Source: https://docs.freeplay.ai/api-reference/agent-datasets/update-agent-test-case https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/{test_case_id} Modify an existing agent test case's inputs, outputs, or metadata. Use to correct errors or update expected outputs. # Create Agent Source: https://docs.freeplay.ai/api-reference/agents/create-agent https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/agents Create a new agent in your project. Agents are used to organize and group related sessions and traces. # Delete Agent Source: https://docs.freeplay.ai/api-reference/agents/delete-agent https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agents/{agent_id} Delete an agent from your project. This operation cannot be undone. # Get Agent Source: https://docs.freeplay.ai/api-reference/agents/get-agent https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agents/{agent_id} Retrieve details for a specific agent by ID. # List Agents Source: https://docs.freeplay.ai/api-reference/agents/list-agents https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agents Retrieve a paginated list of agents for your project. Optionally filter by name. `page_size` defaults to 30, maximum 100. # Update Agent Source: https://docs.freeplay.ai/api-reference/agents/update-agent https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/agents/{agent_id} Update the name of an existing agent. # Add Project Member Source: https://docs.freeplay.ai/api-reference/configuration/add-project-member https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/members Grant a user access to the project with a specified role. Requires project admin role. # Create Environment Source: https://docs.freeplay.ai/api-reference/configuration/create-environment https://app.freeplay.ai/openapi.json post /api/v2/environments Create a new deployment environment. Use to add environments beyond Freeplay defaults, like staging or feature branches. # Create Project Source: https://docs.freeplay.ai/api-reference/configuration/create-project https://app.freeplay.ai/openapi.json post /api/v2/projects Create a new project in your workspace. Use to organize prompts, datasets, and observability data by team or application. # Create User Source: https://docs.freeplay.ai/api-reference/configuration/create-user https://app.freeplay.ai/openapi.json post /api/v2/users Create a new user in your workspace. Requires account admin role. Use for automated user provisioning or SCIM integrations. # Delete Environment Source: https://docs.freeplay.ai/api-reference/configuration/delete-environment https://app.freeplay.ai/openapi.json delete /api/v2/environments/{environment_id} Remove an environment. Use when retiring deployment targets no longer in use. # Delete Project Source: https://docs.freeplay.ai/api-reference/configuration/delete-project https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id} Permanently delete a project and all its associated data. Requires project admin role. # Delete User Source: https://docs.freeplay.ai/api-reference/configuration/delete-user https://app.freeplay.ai/openapi.json delete /api/v2/users/{user_id} Remove a user from your workspace. Requires account admin role. Use for offboarding or access revocation. # Get Project Source: https://docs.freeplay.ai/api-reference/configuration/get-project https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id} Retrieve the current project's details. Use to check project settings like spend limits or data retention. # Get User Source: https://docs.freeplay.ai/api-reference/configuration/get-user https://app.freeplay.ai/openapi.json get /api/v2/users/{user_id} Retrieve a user's details by ID. Requires account admin role. Use to check user settings or role assignments. # List Environments Source: https://docs.freeplay.ai/api-reference/configuration/list-environments https://app.freeplay.ai/openapi.json get /api/v2/environments Retrieve all deployment environments in your workspace. Use to discover available environments for prompt deployment. `page_size` defaults to 30, maximum 100. # List Project Members Source: https://docs.freeplay.ai/api-reference/configuration/list-project-members https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/members Retrieve all users with access to the project and their roles. Use for access management or auditing. `page_size` defaults to 30, maximum 100. # List Projects Source: https://docs.freeplay.ai/api-reference/configuration/list-projects https://app.freeplay.ai/openapi.json get /api/v2/projects/all Retrieve all projects accessible to the current user. Use to discover available projects or build project selection UIs. # List Users Source: https://docs.freeplay.ai/api-reference/configuration/list-users https://app.freeplay.ai/openapi.json get /api/v2/users Retrieve all users in your workspace. Requires account admin role. Use for user management or access auditing. Set include_deleted=true to include soft-deleted users. # Remove Project Member Source: https://docs.freeplay.ai/api-reference/configuration/remove-project-member https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/members/{user_id} Revoke a user's access to the project. Requires project admin role. # Update Environment Source: https://docs.freeplay.ai/api-reference/configuration/update-environment https://app.freeplay.ai/openapi.json patch /api/v2/environments/{environment_id} Rename an existing environment. Use when consolidating or reorganizing deployment targets. # Update Project Source: https://docs.freeplay.ai/api-reference/configuration/update-project https://app.freeplay.ai/openapi.json put /api/v2/projects/{project_id} Modify project settings like name, visibility, or resource limits. Requires project admin role. # Update Project Member Source: https://docs.freeplay.ai/api-reference/configuration/update-project-member https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/members/{user_id} Change a user's role in the project. Requires project admin role. # Update User Source: https://docs.freeplay.ai/api-reference/configuration/update-user https://app.freeplay.ai/openapi.json patch /api/v2/users/{user_id} Modify a user's name, role, or profile settings. Requires account admin role. # Create Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-code-evaluation-criteria https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/code Create a new code evaluation criteria with an initial auto-deployed version. # Create Code Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-code-evaluation-criteria-version https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions Create a new version of an existing code evaluation criteria. The version is not deployed until published. # Create Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-evaluation-criteria https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria Create a new evaluation criteria in a project. Use to set up criteria for evaluating LLM outputs in test runs or online evaluations. # Create Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-evaluation-criteria-version https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions Create a new version of an evaluation criteria. Use to iterate on evaluation criteria configuration while preserving the version history. # Delete Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/delete-code-evaluation-criteria https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id} Permanently delete a code evaluation criteria and all its versions. # Delete Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/delete-evaluation-criteria https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id} Permanently delete an evaluation criteria and all its versions. Use when retiring criteria no longer needed. # Delete Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/delete-evaluation-criteria-version https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions/{version_id} Permanently delete an evaluation criteria version. Use to clean up draft or unused versions. # Deploy Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/deploy-evaluation-criteria-version https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions/{version_id}/deploy Activate an evaluation criteria version for use in evaluations. Use to promote a tested version to production. # Disable Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/disable-code-evaluation-criteria https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/disable Deactivate a code evaluation criteria without deleting it. # Disable Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/disable-criteria https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/disable Deactivate an evaluation criteria without deleting it. Use to temporarily pause the use of a given criteria in evaluations. # Enable Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/enable-code-evaluation-criteria https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/enable Activate a code evaluation criteria for use in evaluations. # Enable Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/enable-criteria https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/enable Activate an evaluation criteria for use in test runs and online evaluations. # Execute Evals for Completion Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/execute-evals-for-completion https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/completions/{completion_id}/evaluations Queue evaluations to run against a specific completion. Evaluations are executed asynchronously. Use the `eval_types` field to control which evaluations run: - `all` (default): Run both LLM-as-judge and code evaluations - `llm-as-judge`: Run only LLM-based evaluations - `code`: Run only code-based evaluations Optionally pass `criteria_ids` to run specific evaluation criteria. When provided, those criteria are executed regardless of whether they have already been evaluated. When omitted, only criteria that have not yet been evaluated for this completion will run. # Execute Evals for Trace Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/execute-evals-for-trace https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/traces/{trace_id}/evaluations Queue evaluations to run against a specific trace. Evaluations are executed asynchronously. Use the `eval_types` field to control which evaluations run: - `all` (default): Run both LLM-as-judge and code evaluations - `llm-as-judge`: Run only LLM-based evaluations - `code`: Run only code-based evaluations Optionally pass `criteria_ids` to run specific evaluation criteria. When provided, those criteria are executed regardless of whether they have already been evaluated. When omitted, only criteria that have not yet been evaluated for this trace will run. # Get Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-code-evaluation-criteria https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id} Retrieve a code evaluation criteria and its latest version, including eval code. # Get Code Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-code-evaluation-criteria-version https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions/{version_id} Retrieve a specific version of a code evaluation criteria. # Get Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-evaluation-criteria https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id} Retrieve an evaluation criteria's configuration by ID. Use to inspect criteria settings or get the latest version ID. # Get Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-evaluation-criteria-version https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions/{version_id} Retrieve a specific evaluation criteria version. Use to inspect version configuration or compare with other versions. # List Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-code-evaluation-criteria https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code Retrieve all code evaluation criteria in a project. Optionally filter by target type and ID. `page_size` defaults to 30, maximum 100. # List Code Evaluation Criteria Versions Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-code-evaluation-criteria-versions https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions Retrieve all versions of a code evaluation criteria. `page_size` defaults to 30, maximum 100. # List Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-evaluation-criteria https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria Retrieve all evaluation criteria in a project. Use to discover available criteria or build criteria management UIs. `page_size` defaults to 30, maximum 100. # List Evaluation Criteria Versions Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-evaluation-criteria-versions https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions Retrieve all versions of an evaluation criteria. Use to view version history or compare changes over time. `page_size` defaults to 30, maximum 100. # Publish Code Evaluation Criteria Version Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/publish-code-evaluation-criteria-version https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions/{version_id}/publish Deploy a version, undeploying the currently deployed version. # Reorder Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/reorder-criteria https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/reorder Change the display order of evaluation criteria in the Freeplay UI. Use to organize criteria in a logical sequence for human review workflows. # Update Code Evaluation Criteria Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/update-code-evaluation-criteria https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id} Update criteria-level metadata. Does not create a new version. # Get Insights Source: https://docs.freeplay.ai/api-reference/insights/get-insights https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/insights Retrieve a paginated list of insights for a project. Use to discover insights or filter by prompt template or agent. `page_size` defaults to 30, maximum 100. # Create Model Source: https://docs.freeplay.ai/api-reference/models/create-model https://app.freeplay.ai/openapi.json post /api/v2/models Create or upsert a custom model configuration for the account. # Delete Model Source: https://docs.freeplay.ai/api-reference/models/delete-model https://app.freeplay.ai/openapi.json delete /api/v2/models/{model_id} Delete a custom model configuration from the account. # Get Model Source: https://docs.freeplay.ai/api-reference/models/get-model https://app.freeplay.ai/openapi.json get /api/v2/models/{model_id} Retrieve details for a specific custom model by ID. # List Models Source: https://docs.freeplay.ai/api-reference/models/list-models https://app.freeplay.ai/openapi.json get /api/v2/models Retrieve a paginated list of custom models configured for the account. `page_size` defaults to 30, maximum 100. # Update Model Source: https://docs.freeplay.ai/api-reference/models/update-model https://app.freeplay.ai/openapi.json put /api/v2/models/{model_id} Update an existing custom model configuration. # Add Completion Feedback Source: https://docs.freeplay.ai/api-reference/observability/add-completion-feedback https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completion-feedback/id/{completion_id} Record end-user feedback on a completion. Use to capture thumbs up/down ratings or custom feedback attributes. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/customer-feedback) The `freeplay_feedback` field must be "positive" or "negative". Additional custom fields are supported. # Add Trace Feedback Source: https://docs.freeplay.ai/api-reference/observability/add-trace-feedback https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/trace-feedback/id/{trace_id} Record end-user feedback on a trace. Use to capture feedback on conversation turns or agent workflow outcomes. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/customer-feedback) The `freeplay_feedback` field must be "positive" or "negative". Additional custom fields are supported. # Delete Session Source: https://docs.freeplay.ai/api-reference/observability/delete-session https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/sessions/{session_id} Permanently delete a session and all associated completions. Use when removing test data or honoring data deletion requests. # Record Completion Source: https://docs.freeplay.ai/api-reference/observability/record-completion https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/completions Log an LLM completion with its prompt, response, and metadata. This is the primary endpoint for observability. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/recording-completions) Sessions are created implicitly—just generate a session_id client-side (UUID v4). Optionally provide your own completion_id too to avoid waiting for Freeplay's response. # Record Trace Source: https://docs.freeplay.ai/api-reference/observability/record-trace https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/traces/id/{trace_id} Create or update a trace within a session. Use to group related completions in agent workflows. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/traces) Generate the trace_id client-side (UUID v4) to avoid round-trip latency. # Update Completion Source: https://docs.freeplay.ai/api-reference/observability/update-completion https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completions/{completion_id} Append messages or evaluation results to an existing completion. Use for streaming responses or adding post-hoc evaluation metrics. # Update Session Metadata Source: https://docs.freeplay.ai/api-reference/observability/update-session-metadata https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/id/{session_id}/metadata Merge custom metadata into an existing session. Use to enrich sessions with post-hoc context like user ID or business metrics. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/sessions) # Update Trace by ID Source: https://docs.freeplay.ai/api-reference/observability/update-trace-by-id https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/{session_id}/traces/id/{trace_id} Update a trace's output, metadata, feedback, eval results, and/or test run info by its trace ID. # Update Trace by OTEL Span ID Source: https://docs.freeplay.ai/api-reference/observability/update-trace-by-otel-span-id https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/{session_id}/traces/otel-span-id/{otel_span_id_hex} Update a trace's output, metadata, feedback, eval results, and/or test run info by its OpenTelemetry span ID (hex string). Note: OTEL spans are mapped to Freeplay Traces, so we use the OTEL span ID to identify Traces. # Update Trace Metadata Source: https://docs.freeplay.ai/api-reference/observability/update-trace-metadata https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/{session_id}/traces/id/{trace_id}/metadata Merge custom metadata into an existing trace. Use to enrich traces with post-hoc context. # Bulk Create Prompt Test Cases Source: https://docs.freeplay.ai/api-reference/prompt-datasets/bulk-create-prompt-test-cases https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/bulk Add multiple test cases to a dataset in a single request. Use for batch imports, e.g. from CSV or production logs. Maximum 100 test cases per request. # Bulk Delete Prompt Test Cases Source: https://docs.freeplay.ai/api-reference/prompt-datasets/bulk-delete-prompt-test-cases https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/bulk Remove multiple test cases in a single request. Use for batch cleanup operations. Maximum 100 test cases per request. # Create Prompt-Level Dataset Source: https://docs.freeplay.ai/api-reference/prompt-datasets/create-prompt-level-dataset https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-datasets Create a new dataset for prompt-level testing. Use to organize test cases for evaluating individual prompts. # Delete Prompt Dataset Source: https://docs.freeplay.ai/api-reference/prompt-datasets/delete-prompt-dataset https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-datasets/{dataset_id} Archive a prompt dataset and its test cases. Use when retiring datasets no longer needed. # Delete Prompt Test Case Source: https://docs.freeplay.ai/api-reference/prompt-datasets/delete-prompt-test-case https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/{test_case_id} Remove a test case from a dataset. Use to clean up invalid or outdated test cases. # Get Prompt Dataset Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-dataset https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/id/{dataset_id} # Get Prompt Dataset Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-dataset-1 https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets/{dataset_id} Retrieve a prompt dataset's metadata by ID. Use to check dataset configuration or input schema. # Get Prompt Dataset by Name Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-dataset-by-name https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/name/{dataset_name} # Get Prompt Test Case Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-test-case https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/{test_case_id} Retrieve a specific test case (dataset row) by ID. Use to inspect test case details or debug test run results. # List Prompt Datasets Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-datasets https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets Retrieve all prompt-level datasets in a project. Use to discover available datasets for test runs. `page_size` defaults to 30, maximum 100. # List Prompt Test Cases Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-test-cases https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/id/{dataset_id}/test-cases `page_size` defaults to 10, maximum 10. # List Prompt Test Cases Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-test-cases-1 https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases Retrieve all row-level test cases in a prompt dataset. Use to review dataset contents or export for analysis. `page_size` defaults to 30, maximum 100. # List Prompt Test Cases by Dataset Name Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-test-cases-by-dataset-name https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/name/{dataset_name}/test-cases `page_size` defaults to 10, maximum 10. # Update Prompt Dataset Source: https://docs.freeplay.ai/api-reference/prompt-datasets/update-prompt-dataset https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/prompt-datasets/{dataset_id} Modify a prompt dataset's name, description, or input schema. Use to evolve datasets as prompt requirements change. # Update Prompt Test Case Source: https://docs.freeplay.ai/api-reference/prompt-datasets/update-prompt-test-case https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/{test_case_id} Modify an existing test case's inputs, output, or metadata. Use to change ground truth output values or correct errors. # Upload Prompt Test Cases Source: https://docs.freeplay.ai/api-reference/prompt-datasets/upload-prompt-test-cases https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/datasets/id/{dataset_id}/test-cases Add test cases to a dataset using the legacy upload format. Use for bulk imports with input/output pairs. Maximum 100 examples per request. # Cancel a pending or in-progress prompt optimization job. Source: https://docs.freeplay.ai/api-reference/prompt-optimization/cancel-a-pending-or-in-progress-prompt-optimization-job https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-optimization-jobs/{job_id}/cancel Only jobs that have not yet completed can be cancelled. # Get prompt optimization job status and details. Source: https://docs.freeplay.ai/api-reference/prompt-optimization/get-prompt-optimization-job-status-and-details https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-optimization-jobs/{job_id} Returns the current status of the prompt optimization job, including progress information for polling. When complete, includes the optimized_version_id and optionally test_run_id and comparison_id if run_test_after_optimization was True. # List prompt optimization jobs for the project. Source: https://docs.freeplay.ai/api-reference/prompt-optimization/list-prompt-optimization-jobs-for-the-project https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-optimization-jobs Returns a paginated list of prompt optimization jobs, sorted by creation date (newest first). Optionally filter by status. `page_size` defaults to 30, maximum 100. # Start a new prompt optimization job. Source: https://docs.freeplay.ai/api-reference/prompt-optimization/start-a-new-prompt-optimization-job https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-optimization-jobs Creates an asynchronous job that optimizes the specified prompt template version. The job will analyze examples from the dataset and generate an improved prompt. When `run_test_after_optimization` is True (default), the job will also run baseline and optimized test runs and create a comparison. Poll GET /prompt-optimization-jobs/{job_id} to check job status. # Create prompt template Source: https://docs.freeplay.ai/api-reference/prompt-templates/create-prompt-template https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates Create a new prompt template without any versions. Use when you need to reserve a template name before adding versions. # Create prompt template version by ID Source: https://docs.freeplay.ai/api-reference/prompt-templates/create-prompt-template-version-by-id https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions Create a new version of an existing prompt template. Freeplay assigns a random ID. Use when you have the template ID from a previous API call. **Version Creation Semantics:** A new version is created only if there is no existing version with identical: - Content (prompt messages) - Model - LLM parameters - Version name (if provided) - Version description - Tool schema - Output schema If an identical version exists in any of the target environments, that version is reused and its environments are updated to match the requested environments. This ensures you never create duplicate versions with the same configuration. **Environment Deployment:** - When creating a new version, it will be deployed to the specified environments (or "latest" if none specified) - When reuploading an existing prompt template with a new environment, the environments on that prompt template version will be updated These checks are performed so that you can safely upload prompt templates without worrying that you will be creating duplicate versions. This is especially useful in development workflows where your prompts are stored in code. # Create prompt template version by name Source: https://docs.freeplay.ai/api-reference/prompt-templates/create-prompt-template-version-by-name https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/name/{template_name}/versions Create a new version of a prompt template, referenced by a string or name you define. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/prompts) Set `create_template_if_not_exists=true` to auto-create the template if it doesn't exist. **Version Creation Semantics:** A new version is created only if there is no existing version with identical: - Content (prompt messages) - Model - LLM parameters - Version name (if provided) - Version description - Tool schema - Output schema If an identical version exists in any of the target environments, that version is reused and its environments are updated to match the requested environments. This ensures you never create duplicate versions with the same configuration. **Environment Deployment:** - When creating a new version, it will be deployed to the specified environments (or "latest" if none specified) - When reuploading an existing prompt template with a new environment, the environments on that prompt template version will be updated These checks are performed so that you can safely upload prompt templates without worrying that you will be creating duplicate versions. This is especially useful in development workflows where your prompts are stored in code. # Delete prompt template Source: https://docs.freeplay.ai/api-reference/prompt-templates/delete-prompt-template https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-templates/id/{template_id} Archive a prompt template and all its versions. Use when retiring templates that are no longer needed. This is a soft delete. # Delete prompt template version Source: https://docs.freeplay.ai/api-reference/prompt-templates/delete-prompt-template-version https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id} # Get all prompt templates by environment Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-all-prompt-templates-by-environment https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/all/{environment} # Get bound prompt template by name and environment Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-bound-prompt-template-by-name-and-environment https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/name/{name} # Get bound prompt template version by ID Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-bound-prompt-template-version-by-id https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id} # Get prompt template Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-prompt-template https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/id/{template_id} Retrieve a prompt template's metadata by ID. Use to check if a template exists or get its latest version ID. # Get prompt template by name and environment Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-prompt-template-by-name-and-environment https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/name/{name} # Get prompt template version by ID Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-prompt-template-version-by-id https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id} # List prompt template versions Source: https://docs.freeplay.ai/api-reference/prompt-templates/list-prompt-template-versions https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions Retrieve all versions of a prompt template. Use to view version history or compare changes over time. `page_size` defaults to 30, maximum 100. # List prompt templates Source: https://docs.freeplay.ai/api-reference/prompt-templates/list-prompt-templates https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates Retrieve all prompt templates in a project. Use to discover available templates or build template management UIs. `page_size` defaults to 30, maximum 100. # Update environment for prompt template version Source: https://docs.freeplay.ai/api-reference/prompt-templates/update-environment-for-prompt-template-version https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id}/environments Deploy a prompt version to one or more environments. Use to promote versions through dev, staging, and production. # Update prompt template Source: https://docs.freeplay.ai/api-reference/prompt-templates/update-prompt-template https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/prompt-templates/id/{template_id} Rename a prompt template. Use when refactoring template names across your codebase. # List Review Queues Source: https://docs.freeplay.ai/api-reference/review-queues/list-review-queues https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/review-queues Retrieve a paginated list of review queues for the project, including creator, assignees, and review progress counts. `page_size` defaults to 30, maximum 100. # Create Saved Search Source: https://docs.freeplay.ai/api-reference/saved-searches/create-saved-search https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/session-filters Create or upsert a saved search for the project. # Delete Saved Search Source: https://docs.freeplay.ai/api-reference/saved-searches/delete-saved-search https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/session-filters/{filter_id} Delete a saved search from the project. # Get Saved Search Source: https://docs.freeplay.ai/api-reference/saved-searches/get-saved-search https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/session-filters/{filter_id} Retrieve details for a specific saved search by ID. # List Saved Searches Source: https://docs.freeplay.ai/api-reference/saved-searches/list-saved-searches https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/session-filters Retrieve a paginated list of saved searches for the project. `page_size` defaults to 30, maximum 100. # Update Saved Search Source: https://docs.freeplay.ai/api-reference/saved-searches/update-saved-search https://app.freeplay.ai/openapi.json put /api/v2/projects/{project_id}/session-filters/{filter_id} Update an existing saved search. # Get All Completion Statistics Source: https://docs.freeplay.ai/api-reference/search-&-analytics/get-all-completion-statistics https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completions/statistics Retrieve aggregate evaluation statistics across all prompts for a date range. Use for dashboard metrics or trend analysis. Maximum date range is 30 days. # Get Completion Statistics Source: https://docs.freeplay.ai/api-reference/search-&-analytics/get-completion-statistics https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completions/statistics/{prompt_template_id} Retrieve evaluation statistics for a specific prompt template. Use to track quality metrics for individual prompts. Maximum date range is 30 days. # List Sessions Source: https://docs.freeplay.ai/api-reference/search-&-analytics/list-sessions https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/sessions Retrieve sessions with their completions, ordered by most recent first. Use to display conversation history or analyze LLM usage patterns. Traces are referenced by ID but not expanded. Filter by custom metadata using query parameters prefixed with `custom_metadata.` (e.g., `custom_metadata.user_id=123`). `page_size` defaults to 10, maximum 100. Prefer using the `/search/sessions` endpoint for more advanced filtering. # Search Completions Source: https://docs.freeplay.ai/api-reference/search-&-analytics/search-completions https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/search/completions Query LLM completions using advanced filters. Use to find specific prompts and responses, filter by evaluation results or metadata, prompt templates, latency, and more. Supports pagination and optional inclusion of all child traces and completions within the session, using the `include_children` parameter. `page_size` defaults to 30, maximum 100. For filter operators and examples, see: https://docs.freeplay.ai/openapi/search-api-operators # Search Sessions Source: https://docs.freeplay.ai/api-reference/search-&-analytics/search-sessions https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/search/sessions Query sessions using advanced filters. Use to find specific conversations or filter by metadata. Supports pagination and optional inclusion of all traces and completions within the session, using the `include_children` parameter. `page_size` defaults to 30, maximum 100. For filter operators and examples, see: https://docs.freeplay.ai/openapi/search-api-operators # Search Traces Source: https://docs.freeplay.ai/api-reference/search-&-analytics/search-traces https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/search/traces Query traces using advanced filters. Use to find specific trace executions, filter by metadata, or analyze trace patterns. Supports pagination and optional inclusion of all child traces and completions within the session, using the `include_children` parameter. `page_size` defaults to 30, maximum 100. For filter operators and examples, see: https://docs.freeplay.ai/openapi/search-api-operators # Create Tag Source: https://docs.freeplay.ai/api-reference/tags/create-tag https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/tags Create a new tag in the project. If color is not provided, it will be auto-assigned. # List Project Tags Source: https://docs.freeplay.ai/api-reference/tags/list-project-tags https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/tags Retrieve all tags in a project. Optionally filter by entity type (e.g., 'dataset'). # Create Test Run Source: https://docs.freeplay.ai/api-reference/test-runs/create-test-run https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/test-runs # Get Test Run Results Source: https://docs.freeplay.ai/api-reference/test-runs/get-test-run-results https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/test-runs/id/{test_run_id} # List Test Runs Source: https://docs.freeplay.ai/api-reference/test-runs/list-test-runs https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/test-runs `page_size` defaults to 100, maximum 100. # AI Features Source: https://docs.freeplay.ai/core-concepts/ai/ai-features How Freeplay uses AI to accelerate your product improvement workflow. Surface patterns and root causes from your evaluation data, human reviews, and test runs Score individual completions and traces using LLMs to evaluate your AI outputs at scale Create better evaluation criteria with AI-powered suggestions and prompt drafts for LLM judges Classify logs to reveal usage patterns and understand how users interact with your AI Get AI-generated suggestions for improved prompts based on your production data ## Overview All AI features in Freeplay work by calling LLM APIs to analyze your data. They are designed to work with different models and to use your API keys and model preferences, based on your account settings. ## Managing AI feature settings ### Disabling specific features Individual AI features can be controlled through their respective configuration: * **Model-graded evaluations**: Disable per evaluation by turning off or setting sample rate to zero * **Eval Creation Assistant**: This is an on-demand feature that only runs when creating evals * **Auto-categorization**: Disable per auto-category by turning off or setting sample rate to zero * **Prompt optimization**: This is an on-demand feature that only runs when triggered * **Review Insights**: Runs automatically during review; disable via the Insights toggle in Project Settings > AI Features * **Evaluation Insights**: Runs weekly; disable via the Insights toggle in Project Settings > AI Features ### Cost considerations AI features consume tokens from the selected LLM provider. Costs depend on: * Which features you use and how frequently * The volume of data being analyzed * The models being used (more capable models typically cost more) When Freeplay Keys are enabled, Freeplay covers the cost of AI feature usage. When using your own API keys, costs are billed directly to your provider account. Token usage for AI features is tracked separately from your application's LLM usage and is visible in the Usage dashboard. If you're using your own API keys, monitor this usage and consider adjusting feature sampling rates if costs are higher than expected. ## Related pages * [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure LLM judges for automated scoring * [Auto-Categorization](/core-concepts/evaluations/auto-categorization) — Set up automatic content classification * [Review Queues](/core-concepts/review-queues) — Organize human review workflows and surface insights * [Model Management](/account-setup/model-management) — Configure LLM providers and API keys # Overview Source: https://docs.freeplay.ai/core-concepts/ai/ai-insights How Freeplay's AI Insights agent analyzes your data to surface actionable findings and improvement opportunities. AI Insights can be toggled off in **Project Settings > AI Features**. Freeplay's AI Insights is an intelligence layer that sits at each transition point in your agent development workflow, helping you move from raw data to decisions. A background AI agent analyzes your labeled data — evaluation results, human annotations, and test runs — to generate findings. These findings help you understand the **why** behind performance changes and what to do about it. The goal is simple: every time you start spending time in Freeplay, you should quickly be able to spot the next most impactful thing you can do to improve your AI system. Then after every labeling session or experiment you run, you should quickly know what to do next. ## The problem Insights solve You can quickly log hundreds of thousands or millions of traces. You can run LLM judges or other metrics to score those logs, and visualize those in dashboards showing pass/fail rates. Those might tell you that you have a problem, but scores rarely tell you how to fix anything. Freeplay helps you track evaluation performance over time but raw metrics alone stop short of helping you understand **why** performance changed or **what to do about it**. When you try to decide what to actually improve, you're stuck asking the same questions: Why is that metric failing? What should I do to fix it? Where should I start first? AI Insights closes that gap by proactively generating findings that add a layer of interpretation on top of your raw metrics — turning scores into direction. See our [blog post](https://freeplay.ai/blog/automated-insights-for-ai-agents) for more information. ## How Insights work Insights come from Freeplay's own AI agent that analyzes your data from multiple sources and generates actionable findings. ### Where Insights run Insights run in two places across the Freeplay platform, each representing a decision point in the AI quality workflow: | Location | Helps you answer | | :------------------ | :------------------------------------------------------------------------------------------------------------ | | **Production logs** | Where should I focus? What's broken that I didn't know about? | | **Human reviews** | What patterns are emerging across my team's annotations? What are the root causes of issues people have seen? | Each of these represents a moment where you need to interpret lots of data and decide what to do next — exactly the kind of work AI is good at. A single completion or trace can be tagged with more than one insight. ### Types of Insights Freeplay generates two types of Insights, each tied to a different data source and decision point in your workflow: * [**Evaluation Insights -**](/core-concepts/ai/evaluation-insights) Analyze production logs scored by LLM-as-a-judge evaluations. Run on a weekly cadence to surface systemic issues across your logged data. * [**Review Insights -**](/core-concepts/ai/review-insights) Analyze human annotations in real time. Every note, label, or evaluation triggers the agent to identify patterns and group them into themes. ### What data insights use Each type of insight is based on different types of information within Freeplay. Here are the sources for each insight: | | [**Evaluation**](/core-concepts/ai/evaluation-insights) Insights | Review Insights | | ------------------------ | ---------------------------------------------------------------- | --------------- | | Model-graded evaluations | :check | :check | | Human evaluations | X | :check | | Human notes and lables | X | :check | | Logged data | :check | :check | AI Insights does not currently use code evaluations or auto-categorizations as input sources. ## Viewing and using Insights Insights tab showing generated findings with severity levels and linked traces AI Insights are viewable from the **Home page** or the **Insights** tab within your project. ### Refining Insights You can fine-tune insights to improve their accuracy and usefulness: * **Update the name** — providing a more descriptive name can slightly adjust and refine the grouping of tagged records * **Edit the description** — adding more details, specific errors, or patterns found in your analysis helps fine-tune what the insight captures ### Resolving Insights The goal of insights is to **resolve them and surface new ones**. Insights can help lead your team towards solving the key issues in your product. Once an issue is identified and fixed, the insight will start to lose traction as no new information is added to it. ## Insights and the data flywheel Insights provide a clean path towards understanding the **why** behind errors. Some outcomes of insights include: * **Prompt improvements** — actionable suggestions for how to modify your prompts * **New evaluations** — generating new LLM-as-a-judge evals based on discovered patterns * **Deeper investigation** — surfacing issues that people might be missing Combined with [Review Queues](/core-concepts/review-queues), [prompt optimization](/core-concepts/ai/prompt-optimization), and automated evaluations, Insights help your data flywheel operate smoothly. ## Related resources * [AI Features Overview](/core-concepts/ai/ai-features) — Overview of all AI-powered features in Freeplay * [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure the LLM judges that feed into Insights * [Review Queues](/core-concepts/review-queues) — Set up human review workflows that generate Review Insights * [Automated Insights for AI Agents](https://freeplay.ai/blog/automated-insights-for-ai-agents) — Blog post with more detail on the vision behind Insights # Auto-Categorization Source: https://docs.freeplay.ai/core-concepts/ai/auto-categorization Use AI to automatically tag and classify your production logs based on categories you define. Auto-categorization uses AI to automatically tag and classify your production logs based on categories you define. This adds a layer of intelligence that helps you understand usage patterns and identify trends. ## How it works 1. You define category types that align with your business needs (e.g., product areas, user intent types, issue categories) 2. For each category, you provide a clear name and description 3. As logs flow through Freeplay, the AI classifies them according to your categories 4. Categories appear in the observability dashboard for filtering and analysis ## Use cases * **Usage analysis**: Understand what types of questions users ask most frequently * **Issue identification**: Track which product areas generate the most problems * **Dataset curation**: Filter logs by category to build targeted test datasets * **Review queue creation**: Focus review efforts on specific categories ## Configuration Auto-categorization is configured at the prompt template or agent level, similar to other evaluations: 1. Navigate to your prompt template or agent 2. Create a new evaluation with type **Multi-select** 3. Enable auto-categorization and define your categories 4. Each category needs a name (max 32 characters) and description (max 500 characters) 5. Configure whether items can be tagged with multiple categories or just one **Best practice:** Auto-categorization works best with clear, mutually exclusive categories. If you see many items tagged as "Other" or miscategorized, refine your category descriptions. [Learn more about auto-categorization →](/core-concepts/evaluations/auto-categorization) # Eval Creation Assistant Source: https://docs.freeplay.ai/core-concepts/ai/eval-creation-assistant Create better evaluation criteria with AI-powered suggestions and prompt drafts for LLM judges. Writing effective evaluation prompts can be challenging, especially for teams new to LLM-based quality assessment. Freeplay's Eval Creation Assistant uses AI to help you draft better evals faster—whether you're starting from scratch or adapting a template. ## How it works The Eval Creation Assistant helps in two ways: **Create custom evals from scratch**: Start with the basic question you want to answer about your AI's output. The assistant will: 1. Help you refine your evaluation question to be clear and measurable 2. Suggest improvements to your eval structure 3. Automatically draft a model-graded eval prompt tailored to your specific prompts and data **Adapt from templates**: Choose from common evaluation templates like Answer Faithfulness (for RAG), Similarity, Toxicity, or Tone. The assistant will: 1. Automatically customize the template to match your prompt structure 2. Reference the correct input variables from your prompts 3. Generate a ready-to-use eval prompt with one click Because Freeplay knows your prompt structure and has access to real-world examples from your logs, the assistant can generate eval prompts that are specific to your context rather than generic templates. ## Use cases * **Getting started quickly**: Teams new to evals can create their first evaluations without prior experience * **Adopting best practices**: Start with industry-standard eval patterns and customize them for your needs * **Cross-functional collaboration**: Product managers, analysts, and domain experts can contribute to eval creation without writing code ## Using the assistant 1. Navigate to your prompt template or agent 2. Go to the **Evaluations** section 3. Choose **Create your own** or select from the template library 4. For custom evals: Enter your evaluation question and follow the AI's suggestions 5. For templates: Select a template and the AI will automatically adapt it to your prompt 6. Test the generated eval against sample data 7. Use the [alignment flow](/practical-guides/creating-and-aligning-model-graded-evals) to validate that the eval matches human judgment Even when using templates, the AI adapts them to your specific prompt variables and data structure—so you get truly customized evals, not just generic prompts. **Best practice:** If you're new to writing evals or unsure where to start, use the Eval Creation Assistant's "Create your own" option. Describe what you want to evaluate in plain language, and the AI will generate a custom eval prompt tailored to your specific prompts and use case. # Evaluation Insights Source: https://docs.freeplay.ai/core-concepts/ai/evaluation-insights AI-powered analysis of your production evaluation data to surface issues and improvement opportunities. Evaluation Insights analyze your production log data to surface issues you might not catch from dashboards alone. Image ## How they work Freeplay proactively analyzes logged data that has [model-graded evaluations](/core-concepts/evaluations/model-graded-evaluations) applied. The agent reviews these logs and identifies key patterns across the data. Here is the general flow: 1. Freeplay collects evaluation results over a time period (requiring at least 10 logs with evaluation data) 2. The AI analyzes the logs, looking for patterns in: * Poor-scoring outputs and their common characteristics * Correlation between different evaluation criteria * Input patterns that tend to produce poor results 3. The agent then reviews these results, assigns, creates or updates existing insights to properly assign and group the data For each insight, you get a clear description of the problem, an easy link to the underlying traces that back it up, and the number of matching records as a proxy for scale and impact to help you prioritize what matters most. ## When they run Evaluation Insights run on a **weekly cadence**, generating findings every Monday morning based on the previous week's data. Evaluation Insights can be disabled in **Project Settings > AI Features**. ## Related resources * [AI Insights Overview](/core-concepts/ai/ai-insights) — How Freeplay's AI Insights work across the platform * [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure the LLM judges that feed into Insights # Model-Graded Evaluations Source: https://docs.freeplay.ai/core-concepts/ai/model-graded-evaluations Use AI to automatically score your AI outputs based on evaluation criteria you define. Model-graded evaluations (also called LLM judges or auto-evaluations) use AI to automatically score your AI outputs based on criteria you define. This is the foundation of automated quality assessment in Freeplay. ## How it works When you configure a model-graded evaluation: 1. You define the evaluation criteria with a name, question, and scoring type (Yes/No, 1-5 scale, etc.) 2. You write instructions explaining what the LLM should evaluate and provide a rubric with scoring guidelines 3. Freeplay generates a structured prompt that includes your criteria, the completion being evaluated, and relevant context The LLM then scores each completion according to your rubric and provides an explanation for its decision. ## Use cases * **Production monitoring**: Automatically sample and evaluate a subset of production traffic * **Batch testing**: Run evaluations across entire datasets during test runs * **Quality gates**: Identify outputs that fail specific quality thresholds ## Configuration Model-graded evaluations are configured at the prompt template or agent level: 1. Select the **Evaluations** tab from the menu and then select **New Evaluation** 2. Set the target to your **prompt/agent** to evaluate and the type to **Model-graded** 3. Create your own or select from a pre-configured example 4. Write instructions that reference your prompt variables (e.g., `{{inputs.context}}`, `{{output}}`) 5. Define a rubric that maps scores to specific behaviors Use Freeplay's alignment tools to compare auto-evaluation scores against human labels and iteratively improve your evaluation prompts. **Best practice:** Model-graded evaluations are the foundation for many other AI features. Prompt optimization and Evaluation Insights both work better when you have well-configured evaluations generating data. Start here before enabling other AI features. [Learn more about model-graded evaluations →](/core-concepts/evaluations/model-graded-evaluations) # Prompt Optimization Source: https://docs.freeplay.ai/core-concepts/ai/prompt-optimization Use AI to analyze your production data, evaluation results, and customer feedback to suggest improved prompts. Prompt optimization uses AI to analyze your production data, evaluation results, and customer feedback to suggest improved prompts. It can also help update prompts when switching between models. ## How it works 1. You select a prompt template version to optimize and choose a dataset or set of evaluated sessions 2. You configure what data sources to use: * **Human labels**: Scores and feedback from your team's reviews * **Customer feedback**: Direct feedback captured from end users * **Best practices**: Provider-specific prompting guides (OpenAI or Anthropic) 3. You can optionally provide specific instructions about what to improve 4. Freeplay's AI analyzes the data and generates: * An optimized prompt template * An explanation of changes made * A description of the new version ## Use cases * **Prompt iteration**: Get AI-suggested improvements based on where your current prompt is failing * **Model migration**: Update prompts optimized for one model to work well with another * **Data-driven improvement**: Use production signals to guide prompt changes ## Configuration Prompt optimization is available from the prompt template editor: 1. Open a prompt template and select a version 2. Click **Optimize** to open the optimization panel 3. Select your data source (dataset or evaluated sessions) 4. Choose which signals to include (labels, feedback, best practices) 5. Optionally add specific instructions 6. Run the optimization After optimization completes, Freeplay creates a new prompt version and automatically runs a comparative test so you can evaluate the results side-by-side. Prompt optimization works best with at least 10-20 evaluated examples that include a mix of good and poor outputs. # Review Insights Source: https://docs.freeplay.ai/core-concepts/ai/review-insights Review Insights work alongside your human reviewers to perform real-time root cause analysis. As your team reviews completions and traces, the insights agent surfaces patterns and groups them into actionable findings. Review Insights panel showing identified patterns across reviewed completions ## How they work Review Insights workflow showing how human labels feed into AI analysis When humans label data in Freeplay the Insights agent analyzes each reviewed item in the background. It identifies common patterns, groups related items into **themes**, and suggests actions based on what it finds. The inputs to Review Insights include: * **Human labels** — annotations, notes, and scores from human reviewers * **LLM-as-a-judge evaluations** — scores and reasoning from your auto-evaluators applied during review * **Logs** — the completions or traces that were evaluated ## When they run Review Insights run **anytime a human label is added** to data. Every annotation — notes, human evals, or LLM-as-a-judge evals — triggers the agent to analyze and update insights. When combined with [Review Queues](/core-concepts/review-queues) these review insights can point to key issues in your system. Review Insights can be disabled in **Project Settings > AI Features**. Review Insights themes are generated automatically and may occasionally be too broad or too narrow. Regularly review themes and use merge/prune actions to keep them useful. ## Related resources * [AI Insights Overview](/core-concepts/ai/ai-insights) — How Freeplay's AI Insights work across the platform * [Review Queues](/core-concepts/review-queues) — Set up human review workflows that generate Review Insights * [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure the LLM judges that feed into Insights # Curating Useful Datasets for Testing & Evaluation Source: https://docs.freeplay.ai/core-concepts/datasets/dataset-curation Learn strategies for building high-quality datasets that accurately represent real-world usage. ## What are datasets and why are they important? Testing and evaluation are key aspects of the LLM development cycle. Your test quality and reliability are a function of two primary components: your Evaluators and your Datasets. In this guide, we are going to focus on Dataset curation. Simply put, datasets are collections of inputs and outputs that you can use to test your LLM systems. Each example's **output** represents either a golden response (the ideal answer) or a captured failure case from production. Including outputs is strongly recommended — they are what evaluations compare new LLM responses against during test runs. For more details, see [Understanding the Output Field](/core-concepts/datasets/datasets#understanding-the-output-field). Datasets are important because for your tests to be truly informative your datasets need to accurately represent the issues and situations your LLM systems face in the wild. We’ll cover how we at Freeplay think strategically about building datasets and then look tactically at how to curate datasets inside of Freeplay. ## Dataset Curation Strategy Once your team has decided what product or feature you want to build, a common next step is curating your datasets. Many teams will build a single dataset of various scenarios for an LLM feature and instinctively stop there. While one dataset is a fantastic starting point, teams quickly realize that multiple datasets are crucial for effective testing and experimentation. Broadly speaking there are two types of datasets: Targeted datasets and Broad-based datasets. ### Targeted datasets Targeted datasets are datasets that are focused on a narrowly defined issue or situation. For example, let’s say you’re working on an e-commerce use case in which we are using an LLM to answer customer questions about their orders. The pipeline has two components: 1. First, the LLM generates a SQL query from the user question 2. Then, the LLM uses the results of that query to generate an answer To test this pipeline, you might create a targeted dataset called “Query Hallucinations”. This dataset would be a collection of examples in which the model hallucinates a table name. You might create another targeted dataset called “Delivered Orders”, which collects examples where the user asked about an order that was already delivered. The first dataset is focused on a technical failure point and the second dataset is focused on a specific customer situation, orders that have already been delivered. Both datasets are narrowly focused. When iterating on LLMs it’s often useful to focus on one specific problem at a time. Targeted datasets allow you to quickly iterate over your key areas of concern. ### Broad-based datasets Broad-based datasets are datasets that include a wide array of examples and are not focused on any specific issue or situation. The classic example of a broad-based dataset would be the “golden set.” A golden set is a dataset consisting of examples hand-curated by humans to be the ideal output for some given input. Unlike a targeted dataset, broad-based datasets contain a variety of situations meant to capture the totality of cases your LLM system needs to be able to handle. This kind of dataset is often used for benchmarking or regression testing. These two types of datasets can then be used together during experimentation. Continuing our e-commerce example from earlier, let’s say you’re focused on reducing hallucinations in SQL query generation. You can first focus on making prompt and model changes for that specific issue, frequently testing against your targeted dataset along the way. Then once you think you have a fix in hand, you test the new config against your broad-based dataset to ensure you haven’t regressed on other dimensions. ## Anatomy of a Freeplay Dataset Now that you’re familiar with the primary types of datasets, let’s take a look at what a Freeplay dataset consists of. * Name - ex. “Query Hallucinations” * Description - ex. “Sessions where the model referenced an invalid table” * Prompt Compatibility - A set of inputs that your dataset will be compatible with. * Examples - Examples are combinations of inputs and outputs that you as the user save to a dataset. The output should be either a golden response or a captured failure case — see [Understanding the Output Field](/core-concepts/datasets/datasets#understanding-the-output-field) for guidance. ## Curating Datasets in Freeplay ### Step 1: Create a new Dataset Navigate to the Datasets tab and select “Create dataset”. Give your dataset a Name and Description, then set your Prompt Compatibility. You’ll need to decide what prompt(s) you want your dataset to be compatible with. Compatibility is determined by the prompt’s input variables. Datasets can be compatible with multiple prompts as long as those prompts share at least one common input variable. Alternatively, you can create a new Dataset directly from a completion or trace by clicking “+ Dataset” and then hitting the + icon. Note, if you create the dataset this way, prompt compatibility will be intuited for you based on the completion you create the dataset from. ### Step 2: Add Examples There are a number of ways to add examples to a dataset. **Add from completion** From any completion you can hit “Add to dataset” to create a new example from that completion. This is often a big part of the human review flow, as reviewers are labeling data they can actively be building datasets as well. #### Curating Samples from Production Data When adding new samples directly from Observability, you have the option to manually curate them to be representative samples in your dataset. When you open up the "+ Dataset" modal, it will give you the ability to curate any of the inputs, variables, history and output by adding additional messages, tool calls, multimedia and more. This allows you to add quality samples to your dataset making it even more useful for testing different parts of the product such as failure modes, successful cases and more! **Bulk add from observability** You can bulk add completions to a dataset by going to the Observability tab, toggling to the completions view in the table and then selecting the completions you want to add. Often we will see users filter on things like eval values, customer feedback, or other metrics and bulk adding completions from there. **Upload examples** If you have existing examples you can upload them to Freeplay via JSONL. Navigate to the Datasets tab, select your dataset, and click upload. You can read more about formatting the JSONL file [here](/core-concepts/datasets/datasets#uploading-data-using-jsonl). **Manually write examples** You can also write examples directly in the UI. From any dataset click “Create an example” and a form will appear where you can write a new example by hand. You’ll enter values for each input variable as well as the output. It’s okay to leave any of these blank if it makes sense for your example. ### Step 3: Run a Test against your Dataset After you’ve created a dataset you can run a batch test with any of your compatible prompts. Batch tests can be kicked off either from the Freeplay app or via the SDK. To run a batch test from the UI go to the Tests tab and click “Run Test”. From there you can configure the test by selecting the prompt version you want to test and the dataset you want to test with. To run a batch test from the SDK see the docs [here](/freeplay-sdk/test-runs). ### Step 4: Managing Datasets You can manage your dataset on an ongoing basis in the Datasets tab. Here you can add, edit, and delete examples ### Bonus: Use your Dataset in the Playground When editing a prompt in the playground you can pull in examples from your dataset and run them in real time to test your changes. In the prompt editor click the folder icon and load in examples ## Key Takeaways Dataset curation is an often overlooked part of the LLM development cycle. Your testing is only as good as your underlying datasets. Having a rich collection of datasets empowers developers to iterate faster and ultimately deliver better, higher-quality AI features for your customers. Freeplay helps facilitate that dataset curations process in a fully integrated platform. # Datasets Source: https://docs.freeplay.ai/core-concepts/datasets/datasets Build and manage test datasets to power evaluations, test runs, and fine-tuning workflows. Datasets in Freeplay are an essential part of organizing data to test your LLM systems. They can also be used to curate data for human review or fine-tuning. Datasets are the foundation of [Test Runs](/core-concepts/test-runs/test-runs) in Freeplay. **API Reference**: Freeplay supports two types of datasets: * **Prompt Datasets** (for component-level testing) * **Agent Datasets** (for end-to-end testing) Each is automatically built to enforce schemas that maintain compatibility with your separate prompts and agents. A key benefit of using Freeplay to curate Datasets is that it's seamless to save new examples that you observe in real-world testing or production to existing Datasets. This keeps the data fresh and representative of the actual use of your application. Datasets can be created to test LLM systems across a variety of scenarios, such as: * **Golden Set:** For detecting regressions vs. your ideal ground truth * **Failure Cases:** For tracking failures you observe and testing in the future to confirm they are fixed * **Red Teaming:** For managing adversarial test cases and confirming appropriate behavior by your system * **Random Samples:** For representative testing across a distributed set of values Instructions on how to save observed data or upload data are below. ## Understanding the output field Every dataset entry has an **output** field. While not strictly required, we strongly recommend including an output for each example — it plays a central role in evaluations and test runs, and examples without an output have limited utility for testing. There are two primary ways to use the output field: * **Golden output:** The output represents the ideal, correct response for the given inputs. This is common in golden sets and broad-based datasets where you want to benchmark new prompt versions against a curated standard. When used in test runs or in the playground, these outputs can be viewed to see how the newly generated data compares to the ideal output. * **Failure case:** The output captures a real failure observed in production — such as a hallucination, incorrect answer, or off-tone response. This is useful for building targeted datasets that track known issues so you can confirm they are fixed in future prompt versions. The output can come from any source — uploaded files, completions saved from observed logs, or manually written examples. When saving completions from production logs, you have the option to edit the output before saving, which allows you to curate it into a golden response or preserve it as a failure case depending on your testing goals. # Curating Datasets Datasets in Freeplay can be curated in one of two ways: by saving completions that are recorded to Freeplay straight from the Sessions view, or by uploading existing test cases to a Dataset. ## Saving Data from Recorded Sessions While working with recorded Sessions or Traces in Freeplay, if you encounter values that are relevant for future testing, you can save it directly. You will be given the option to curate the inputs and outputs before saving to the dataset. This can be useful if you want to make this sample represent a specific type of data sample such as a golden or failure case. This can be done at the trace or completion view. To do this, simply: * Click `+ Dataset` above the completion/trace view * Optionally, make adjustments to the inputs, history or outputs * Select the relevant dataset(s) * Optionally, click the `+` button to create a new dataset from this menu ### Bulk Add You can also select multiple completions or traces at once and add a large group of completions to a dataset at one time, even across pages. * Select the "Completions" or "Traces" view on Observability (instead of Sessions) * Click the radio buttons in the table for the rows you want ## Adding Metadata to Dataset Entries image.png Metadata can now be added to entries in your datasets, allowing you to store additional information with each entry. To add or edit metadata for a dataset entry: 1. Navigate to a specific dataset entry 2. Click the "Edit" option in the dropdown menu 3. In edit mode, you'll see a dedicated "Metadata" section at the top of the entry 4. Add customizable key-value pairs such as: * Customer identifiers (e.g., "customerId": "2382721") 5. Click "Add Metadata" to create additional fields as needed 6. Click "Save" to store your changes # Uploading Datasets If you have existing data that is relevant to use for testing prompts in Freeplay, you can upload it directly as a JSONL or CSV file. Both formats support the same fields. ### Prompt Template Datasets | Field | Required | Description | | ------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `inputs.` | **Yes** (at least one) | Each input variable must be prefixed with `inputs.` (e.g., `inputs.question`, `inputs.job_info`). These map directly to `{{variable_name}}` in your prompt template. | | `history` | No | A JSON array of previous messages representing the conversation history. See [Tool Calls in History](#tool-calls-in-history) below for supported message formats. | | `output` | Yes (empty ok) | The expected or ideal output for the given inputs — either a golden response or a captured failure case. See [Understanding the Output Field](#understanding-the-output-field). | | `metadata.` | No | Additional information on each data sample. Metadata columns must start with `metadata.` (e.g., `metadata.employee_id`, `metadata.source`). | For more details on variable usage, see our [Advanced Prompt Templating](/core-concepts/prompt-management/advanced-prompt-templating-using-mustache) guide. **Tool Calls in History:** The `history` field supports tool call and tool result messages, allowing you to upload conversation histories that include function calling interactions. Assistant messages can contain `tool_call` content blocks, and user messages can contain `tool_result` content blocks. See the samples below for examples. ## How to Upload