# Cloud Provider IAM Authentication
Source: https://docs.freeplay.ai/account-setup/cloud-provider-iam-auth
Configure IAM-based authentication for Amazon Bedrock and Google Vertex AI models.
By default, Freeplay routes requests to LLM providers using API keys or credentials managed by Freeplay. For customers with compliance, security, or access requirements, Freeplay also supports IAM-based authentication flows that let you use your own cloud provider roles and service accounts.
This page covers IAM authentication setup for **Amazon Bedrock** (via AWS assume role) and **Google Vertex AI** (via GCP service account impersonation). **BYOC** customers skip to the BYOC section below.
***
## Amazon Bedrock
\*\*Note: \*\*BYOC customers, see next section.
When no custom credentials are configured, Freeplay uses its own AWS credentials to call Bedrock models in the Freeplay VPC.
### Using your own AWS role
If you need Freeplay to call models in your own AWS account — for compliance reasons, to access private models, or to maintain full control over credentials — you can configure [AWS assume role authentication](https://docs.aws.amazon.com/STS/latest/APIReference/API_AssumeRole.html).
With assume role auth, Freeplay temporarily assumes an IAM role in your AWS account using a short-lived token. You retain full control and can revoke access at any time.
#### How it works
1. Freeplay authenticates with an internal AWS role
2. That role assumes **your** AWS role, validated by a shared External ID
3. Using the resulting short-lived token, Freeplay calls Bedrock models in your account
The External ID is a shared string you configure both in your AWS trust policy and on the Freeplay **Settings > Models** page. It prevents unauthorized parties from assuming your role.
#### Customer setup
Freeplay strongly recommends creating a dedicated, isolated IAM role specifically for Freeplay to access Bedrock.
**Step 1: Create an IAM role for Freeplay**
Create a new IAM role in your AWS account with a trust policy that allows Freeplay's role to assume your role, validated by an External ID.
Contact your Freeplay account team to obtain the Freeplay role ARN to use as the **Principal** in your trust policy. You will also need to generate a unique External ID string — this same string must be configured both in your trust policy and on the Freeplay **Settings > Models** page.
**Step 2: Attach a permissions policy**
Attach the following permissions policy to the role to grant Freeplay access to invoke Bedrock models:
```json theme={null}
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowBedrockInvoke",
"Effect": "Allow",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream"
],
"Resource": "*"
}
]
}
```
You can scope the `Resource` field to specific model ARNs if you want to restrict which Bedrock models Freeplay can invoke.
**Step 3: Configure in Freeplay**
1. Navigate to **Settings > Models** in Freeplay
2. Under Amazon Bedrock, enter:
* Your IAM role ARN (e.g., `arn:aws:iam:::role/`)
* The External ID you set in the trust policy
3. Mark this authentication method as default for the provider
4. Save your configuration
Freeplay will now use assume role authentication for all Bedrock requests.
### BYOC (Bring Your Own Cloud) setup
For customers running Freeplay in their own VPC via [BYOC deployment](/security-compliance/byoc), the authentication flow is simplified:
1. Freeplay uses an implicit role from [AWS IRSA](https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html) (IAM Roles for Service Accounts) attached to the pod — no Freeplay-managed user or role is involved
2. The IRSA role assumes your configured AWS role to retrieve a short-lived token
Customer setup follows the same steps as above (create a role, attach the permissions policy, configure in Freeplay), with one difference:
* The **Principal** in your trust policy should reference the IRSA role ARN provided during your BYOC onboarding, rather than the standard Freeplay role ARN
* The External ID condition is optional but recommended for BYOC deployments
***
## Google Vertex AI
### Default behavior
When no custom credentials are configured, Freeplay uses its own GCP service account to call Vertex AI models.
### Using your own GCP project
To route Vertex AI requests through your own GCP project, you need to grant Freeplay's service account permission to create access tokens in your project.
#### Customer setup
**Step 1: Grant the Service Account Token Creator role**
1. In the GCP project where your Vertex AI models are hosted, go to **IAM & Admin > IAM**
2. Click **Grant Access** at the top of the page
3. Set the following:
* **Principal**: Contact your Freeplay account team for the service account email to use
* **Role**: Service Account Token Creator
4. Click **Save**
Role changes can take 30 seconds to 2 minutes to propagate in GCP. If you see permission errors immediately after saving, wait a moment and try again.
**Step 2: Configure in Freeplay**
1. Navigate to **Settings > Models** in Freeplay
2. Configure your Vertex AI provider settings with your GCP project details
3. Mark this authentication method as default for the provider
4. Save your configuration
Freeplay will now use service account impersonation to call Vertex AI models in your project.
# Configure LiteLLM Proxy Models in Freeplay
Source: https://docs.freeplay.ai/account-setup/configure-litellm-proxy-models-in-freeplay
Set up LiteLLM Proxy as a custom provider to access multiple models through a unified interface.
* **Model Flexibility**: Easily switch between different LLM providers while using only OpenAI code for all model interactions.
* **Unified Interface**: Use a consistent API format regardless of the underlying model
* **Simplified Management**: Access numerous models through a single integration
* **Custom Models**: Reference custom-deployed LLMs with the same workflow
## Setting Up LiteLLM Proxy in Freeplay
### Step 1: Set up LiteLLM Proxy & Add Models
In this example we will use gpt-3.5-turbo and claude-3-5-sonnet via Anthropic. To get set up with LiteLLM you can see the full docs [here](https://docs.litellm.ai/docs/proxy/docker_quick_start). To start, ensure you have a model config file like the one below:
```yaml yaml theme={null}
model_list:
- model_name: gpt-3.5-turbo # Use this exact name in Freeplay
litellm_params:
model: openai/gpt-3.5-turbo
api_key: os.environ/OPENAI_API_KEY
- model_name: claude-3-5-sonnet # Use this exact name in Freeplay
litellm_params:
model: anthropic/claude-3-5-sonnet-20240219
api_key: os.environ/ANTH_API_KEY
```
**Important**: When configuring models, the name in Freeplay must match the `model_name` in your LiteLLM Proxy configuration file.
### Step 2: Configure LiteLLM Proxy as a Custom Provider
1. Navigate to Settings in your Freeplay account
2. Find "Custom Providers" section
3. Enable "LiteLLM Proxy" provider
### Step 3: Add an API Key
1. Click "Add API Key"
2. Name your API key
3. Enter your LiteLLM Proxy Master key
4. Optionally, mark it as the default key for LiteLLM Proxy
### Step 4: Add Your LiteLLM Proxy Models
1. Select "Add a New Model"
2. Optionally select if the model supports tool use
3. Optionally add a display name for the model
4. Enter link to your LiteLLM API Proxy
Note: Token pricing information is automatically fetched from LiteLLM Proxy so you do not need to provide it.
### Step 5: Using LiteLLM Proxy Models in the Prompt Editor
1. Open the prompt editor in Freeplay
2. In the model selection field, search for "LiteLLM Proxy"
3. Select one of your configured models
## Integrating LiteLLM Proxy with Your Code
The following example shows how to configure and use LiteLLM Proxy with Freeplay in your application. The benefit of using LiteLLM Proxy is you only need to configure your calls to work with OpenAI, LiteLLM Proxy will handle all the formatting:
```python python theme={null}
#######################
## Configure Clients ##
#######################
# Configure the Freeplay Client
fp_client = Freeplay(
freeplay_api_key=API_KEY,
api_base=f"{API_URL}"
)
# Call OpenAI using the LiteLLM url and api_key.
# This handles the routing to your models while keeping the response
# in a standard format.
client = OpenAI(
api_key=userdata.get("LITE_LLM_MASTER_KEY"),
base_url=userdata.get("LITE_LLM_BASE_URL")
)
#####################
## Call and Record
#####################
# Get the prompt from Freeplay
formatted_prompt = fp_client.prompts.get_formatted(
project_id=PROJECT_ID,
template_name=prompt_name,
environment=env,
variables=prompt_vars,
history=history
)
# Call the LLM with the fetched prompt and details
start = time.time()
completion = client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
# Extract data from LiteLLM response
completion_message = completion.choices[0].message
tool_calls = completion_message.tool_calls
text_content = completion_message.content
finish_reason = completion.choices[0].finish_reason
end = time.time()
print("LLM response: ", completion)
# Record to Freeplay
## First, store the message data in a Freeplay format
updated_messages = formatted_prompt.all_messages(completion_message)
## Now, record the data directly to Freeplay
completion_log = fp_client.recordings.create(
RecordPayload(
project_id=PROJECT_ID,
all_messages=updated_messages,
inputs=prompt_vars,
session_version_info=session,
trace_info=trace,
prompt_info=formatted_prompt.prompt_info, # Note: you must pass UsageTokens for the cost calculation to function
call_info=
CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end,
UsageTokens(completion.usage.prompt_tokens, completion.usage.completion_tokens))
)
)
```
## Current Limitations
* **Automatic Cost Calculation:** UsageTokens must be passed for cost calculations to work with LiteLLM Proxy. Also, for self hosted models that depend on time, cost calculation is not currently supported.
* **Auto-Evaluation Compatibility**: Model-graded evals that are configured and run by Freeplay do not currently support LiteLLM Proxy models.
## Additional Resources
* [LiteLLM Proxy Server Documentation](https://docs.litellm.ai/docs/proxy/quick_start)
* [LiteLLM Supported Models](https://docs.litellm.ai/docs/providers)
* [Configure, Test & Deploy a Fallback LLM Provider](/practical-guides/configuring-a-fallback-llm-provider-with-freeplay)
* [Voice-Enabled AI with Pipecat, Twilio, and Freeplay](/practical-guides/build-voice-enabled-ai-applications-with-pipecat-twilio-and-freeplay)
```
```
# Model Management
Source: https://docs.freeplay.ai/account-setup/model-management
Configure LLM providers, manage API keys, and control which models your team can use.
Only Freeplay admin role users have permission to manage model access and keys.
***
## Configuring Model Access
Freeplay has first-class support for calling models from common hosts/providers including OpenAI, Anthropic, Azure OpenAI Service, Amazon Bedrock, Amazon SageMaker, Groq, Baseten and more. These providers can be directly configured in the Freeplay UI, including configuring appropriate endpoints and API keys or other relevant credentials. These models can then be used end-to-end in the Freeplay application, including in our playground UI.
At the same time, **our SDKs allow you to call any model you want** and record the results with Freeplay. The Freeplay application then lets you configure those models as part of your prompt templates and experiments. An SDK example of calling other models is [here](/freeplay-sdk/recording-completions#calling-any-model).
You can control which of these models and providers your team is able to use and deploy on the Models page.
Navigate to **Settings > Models** to configure models for your team.
* You can disable Default models (e.g. if your team doesn't have permission to use a given provider)
* You can add your own models and endpoints to default providers, like OpenAI fine-tuned models or Llama 3 on SageMaker
* You can control configurability on prompt templates for any other custom models you might have logged with Freeplay
***
## API Keys & Credentials
### Bringing Your Own Keys to Freeplay
Freeplay allows customers to store their LLM provider keys with Freeplay so that Freeplay's application can route requests in the interactive Prompt Editor and for UI-driven Tests. LLM provider keys stored with Freeplay are only used for routing LLM requests, and will not be surfaced in the Freeplay dashboard or via the Freeplay API.
### Security Matters
Freeplay uses application-level encryption to encrypt customer LLM provider keys both at rest and in transit. Keys are only decrypted prior to routing requests to customer models. This level of encryption is supplementary to transparent data encryption provided by cloud providers. We follow industry best practices of encryption key management including regular key rotation and audit logging of all key access. More details on security [here](/security-compliance/security-overview).
### Customer Key Best Practices
1. Provide a unique API key with finely-scoped access for use by Freeplay. Freeplay's application only needs access to make requests to the inference endpoints for your provider.
* e.g. for OpenAI, we recommend creating a separate Project with a cost limit, and creating a key with only access to the `/v1/chat/completions` endpoint.
2. Use Freeplay's UI to rotate your LLM provider keys following your organization's guidelines. As soon as a key is updated in Freeplay's UI, the old value will be destroyed, and the new value will be used for future requests.
3. Monitor usage of your API keys regularly. Freeplay does not set limits on use of customer keys other than those imposed by the LLM providers.
### Configuring Keys in the Dashboard
You can set a default API key for a given provider by clicking on the provider name in the list. This default key will be used for most situations.
For advanced use, you can also link different keys to each endpoint for a given provider, e.g. if you want to use a different key for a fine-tuned model. Set a different key by clicking on the model name.
### Isolating API keys for test runs
Freeplay [test runs](/core-concepts/test-runs/test-runs) execute evaluations with high parallelism to deliver fast results. Without proper key isolation, this parallel traffic can consume rate limits shared with your production application.
To prevent this, create a dedicated API key for Freeplay that is scoped to its own rate limits, separate from your production keys. This ensures test run traffic never competes with your live application for capacity.
If you use the same API key for both Freeplay and your production application, parallel test runs may trigger rate limits (HTTP 429 errors) that affect your production traffic. Freeplay handles 429 responses with exponential backoff, but this does not protect your production application from hitting the same shared limits.
OpenAI supports project-level API keys with independent rate limits:
1. Go to [platform.openai.com](https://platform.openai.com) → **Settings** → **Projects**
2. Click **Create project** (e.g., "Freeplay")
3. Under the new project, go to **API Keys** → **Create a new key**
4. Go to the project's **Limits** tab and set TPM/RPM caps you are comfortable allocating to Freeplay (e.g., 50% of your org limit)
5. In Freeplay, navigate to **Settings** → **Models** and update your OpenAI key to the new project-scoped key
Anthropic supports workspace-level API keys with independent rate limits:
1. Go to [console.anthropic.com](https://console.anthropic.com) → **Settings** → **Workspaces**
2. Create a new workspace (e.g., "Freeplay")
3. In the workspace, go to **API Keys** → **Create a new key**
4. Go to the workspace's **Limits** tab and set caps
5. In Freeplay, navigate to **Settings** → **Models** and update your Anthropic key to the workspace-scoped key
You cannot set custom rate limits on Anthropic's default workspace. You must create a new workspace to configure independent limits.
If your provider is not listed above, the same principle applies: create a separate API key dedicated to Freeplay with its own rate limits or usage scope. This prevents test run traffic from affecting your production application.
Check your provider's documentation for options like projects, workspaces, or service accounts that support independent rate limiting.
***
## Provider-Specific Configuration
For IAM-based authentication with **Amazon Bedrock** or **Google Vertex AI** (using your own AWS roles or GCP service accounts instead of API keys), see [Cloud Provider IAM Authentication](/account-setup/cloud-provider-iam-auth).
## Configuring OpenAI Fine-Tuned Models
You can configure fine-tuned OpenAI models by navigating to **Settings > Models** and choosing **Add fine-tuned model** under the OpenAI Details heading.
Enter the name of your fine-tuned model as provided by OpenAI. You may also optionally enter a more readable display name and specify an associated API key.
### Calling Your Fine-Tuned Model
When creating or editing Prompt Templates you will now see **fine-tuned** as an option in the Model dropdown. Model Version will then populate with all your fine tuned models.
You can call your fine-tuned model from within the prompt editor as well as use your fine-tuned model in the [Freeplay SDK](/freeplay-sdk/recording-completions#record-an-llm-interaction) just as you would any other OpenAI model. Keying the model name, messages and model parameters off of the prompt object.
# Project Setup
Source: https://docs.freeplay.ai/account-setup/project
Create your Freeplay account, generate API keys, and configure models and environments.
## Create your account
Start by signing up for Freeplay at app.freeplay.ai. Once you've created your account, you'll land in your workspace where you can create projects, configure models, and manage your team.
### Custom Freeplay Subdomains
If you have a custom subdomain, it will act as a unique identifier that links your SDK to your Freeplay instance. Here's how to find and use it:
* Your subdomain is part of your Freeplay URL. For example, if your Freeplay URL is `https://acmecorp.freeplay.ai`, then your subdomain is `acmecorp`.
* Ensure you input this subdomain correctly in your SDK configuration. It's crucial for directing your SDK's requests to the right instance.
***
## Generate Freeplay API Keys
Freeplay supports user scoped and project scoped API keys. User based keys are created at the account settings level and inherit all user permissions. Project scoped keys represent service accounts and are scoped to that project specifically.
### To create a User API key:
1. Navigate to Settings > API Access in your Freeplay dashboard
2. Click "Create API Key"
3. Give your key a descriptive name (e.g., "Production" or "Development")
4. Click the copy button to copy your full API key
5. Store it securely—you won't be able to see it again
### To Create a Project API Key
First, you must have a Freeplay project and navigate to the projects settings. Then as an admin you can:
1. Select "Service Accounts"
2. Create a new service account, provide it a name
3. Create an api key for this service account
For more details on API keys and user roles, see our RBAC guide [here](/account-setup/role-based-access-control).
***
## Configure Models
Before you can start building prompts, you'll need to configure which AI models your team can use. Freeplay might include starter credits to help you explore, but we recommend adding your own API keys from providers like OpenAI or Anthropic for ongoing use.
#### Basic setup:
Navigate to Settings > Models to see available providers and models. By default, you'll see common providers like OpenAI, Anthropic, and others. If you've already added API keys for these providers, you're ready to start building prompts.
#### What you can configure:
* Enable or disable specific models and providers
* Add your LLM provider API keys for use in the playground and tests
* Configure custom endpoints (e.g., Azure OpenAI, fine-tuned models)
* Control which models your team can deploy to production
Need more control? Check out our detailed [Model and Key Management guide](/account-setup/model-management) for more information.
***
## Set Up Environments
Freeplay lets you deploy different prompt versions across multiple environments, making it easy to follow a traditional promotion flow from development to production. Default environments:
* latest - Automatically assigned to new prompt versions
* production - Your stable, live version
* sandbox - For testing before production
* dev - Development environment
***
## Freeplay Project ID
Your Project ID is what connects SDK logs to the right project in Freeplay. To find it, just navigate to your project and look at the URL—the long string of characters after /projects/ is your Project ID. For example: `https://app.freeplay.ai/projects//`.
We recommend that you store this ID as an env variable and reference it as `FREEPLAY_PROJECT_ID`.
***
## Project access settings
Projects can be configured as **public** (accessible to all users in your organization) or **private** (accessible only to specific invited members). Private projects are useful for sensitive data that only a subset of team members should access.
To change project access settings, go to **Edit Project > Project access**. For more details on project visibility and user permissions, see [Private vs. Public Projects](/account-setup/role-based-access-control#private-vs-public-projects).
***
With your account configured, API key ready, and models set up, you're ready to start building with Freeplay. Head to the [Quick Start guide](/getting-started/overview) to create your first prompt and run your first test.
# User Roles and Access Controls
Source: https://docs.freeplay.ai/account-setup/role-based-access-control
Manage team permissions with role-based access controls for users and projects.
## Roles
There are 4 primary roles in the Freeplay product.
### Admins
Admins in Freeplay can perform any action, including inviting and managing other users and configuring models available to your team. Use the admin role for those who will be spearheading the adoption of Freeplay within your organization. We recommend having at least 2-3 admins on your account, large organizations may benefit from having additional admins.
### Users
Users can perform most actions in Freeplay but are primarily restricted from taking actions which could have major impacts on the overall account, particularly as it relates to compliance and user management. Use this role for those who are contributing heavily to day to day development.
Users are *restricted* from the following actions:
* Inviting other users
* Managing other users
* Configuring account-level LLM providers and models
* Setting account-level spend limits
### Analysts
Analysts inherit the restrictions of Users but are additionally prohibited from performing sensitive actions like prompt & model deployment or data deletion. Use this role for those who will be reviewing data, running experiments, and evaluating data, but who won't be actively pushing changes to a production system.
In addition to the aforementioned restrictions of Users, Analysts are also *restricted* from the following actions:
* Deploying prompt templates
* Deleting prompt templates
* Deleting production data
* Deleting dataset examples
* Managing API keys (create or delete)
* Managing environments (create or delete)
* Configuring models (at the account level)
* Project creation
### Guests
Guests have the same permissions as Analysts, except that they can only see data for the specific projects they are invited to. (Unlike Analysts who have account-level/site-wide permissions to access any shared projects.) This role is intended for contractors or other collaborators with a narrow scope of focus.
## Role Enforcement & Project-Level Permissions
Each user will be given a role at the account level. This will determine what the user can and can't do by default in any project (i.e. site-wide, default permissions).
However, a user's role can be changed within the context of a specific project. For example, an account level Analyst can be given an Editor role on a specific project. The user will have all the permission associated with an Editor in that specific project, but will maintain their Analyst role at the account level and across other projects.
## Private vs. Public Projects
There are two levels of accessibility for projects. Project privacy is determined at creation, but can also be changed after creation. To change this setting on a project, go to **Edit Project > Project access** and select either **"Anyone at your organization"** or **"🔒 Only specific members"**.
**Public Projects**
Public projects can be accessed by all users on the account with their default user roles. Users do not need to be specifically granted access to the project in order to access it (except Guests, who must be added). We recommend using this project type for any projects that do not contain sensitive data.
**Private Projects**
Private projects can only be accessed by users who have been directly granted access to the project. Each private projects will have 1 or more project admin who determine access for the project. We recommend using this project type for any projects that contain data that only a subset of internal employees or contractors are authorized to see.
## Service Accounts
**Note:** Only account administrators can create, modify, or delete service accounts and their associated API keys.
Service accounts enable project-scoped API key management for production environments. Each project can have multiple service accounts, and each service account can have multiple API keys, allowing you to separate keys across different environments (such as production and testing) or data sources.
To configure service accounts, navigate to your project and select Service Accounts. From there, you can create new service accounts and generate API keys for each one.
# Single Sign-On and SCIM
Source: https://docs.freeplay.ai/account-setup/sso-and-scim
Enterprise authentication options for centralized user management and automated provisioning.
Freeplay offers enterprise-grade authentication options to help organizations manage access at scale. This page covers Single Sign-On (SSO) via SAML and automated user provisioning through SCIM.
SSO and SCIM are available for **Enterprise** tier customers only. Contact [support@freeplay.ai](mailto:support@freeplay.ai) to enable these features for your account.
## Default Authentication Options
By default, Freeplay supports two authentication methods:
* **Email and password** — Standard username/password authentication
* **Google Workspace SSO** — Sign in with your Google account
## Single Sign-On (SSO) via SAML
Enterprise customers can enable SAML-based Single Sign-On to authenticate users through their organization's Identity Provider (IdP). This allows your team to use their existing corporate credentials to access Freeplay.
### Supported Identity Providers
Thanks to our authentication partner WorkOS, Freeplay supports most major Identity Providers including:
* Okta
* Microsoft Entra ID (Azure AD)
* Cisco Duo
* OneLogin
* JumpCloud
* And many others
### Enabling SSO
To enable SSO for your organization:
1. **Contact Freeplay** — Reach out to [support@freeplay.ai](mailto:support@freeplay.ai) to request SSO enablement
2. **Self-serve configuration** — We'll send your IT/auth administrator an email invite to configure SSO through WorkOS
3. **Complete setup** — Your admin completes the SAML configuration in your IdP
4. **Activation** — Once SSO is enabled, other authentication methods (email/password, Google) are disabled for your account
5. **Account creation & role management** — You will continue to add/remove users and update roles manually via the Freeplay UI unless you choose to additioanlly enable SCIM (see below)
Once SSO is enabled, users will only be able to authenticate through your organization's Identity Provider. Make sure your IdP configuration is correct before completing the transition.
## Automated User Provisioning with SCIM
Freeplay offers SCIM (System for Cross-domain Identity Management) support for Enterprise customers to automate the provisioning and deprovisioning of user accounts. This allows you to manage Freeplay user access directly from your Identity Provider (IdP).
### Overview
Enabling SCIM/Directory Sync streamlines your user management by treating your directory as the single source of truth.
* **Automated Onboarding** — New users assigned to Freeplay groups in your IdP are automatically created in Freeplay
* **Automated Offboarding** — Deactivating a user in your IdP immediately revokes their access to Freeplay
* **Centralized Role Management** — User roles are determined solely by their group membership in your directory
### Prerequisites and Setup
To enable SCIM for your account:
1. **Contact Support** — Reach out to the Freeplay team to request SCIM enablement. Please provide the email address of the person who will handle the configuration. This should be somebody that is appropriately permissioned to manage your IdP.
2. **Configuration** — We will send an invite link to configure your directory sync via WorkOS (our third-party auth provider). This link will be in the form of `https://setup.freeplay.ai/init?setupLinkToken=`.
3. **Integration** — Once the configuration is done, either synchronously or async, let us know that you're ready to schedule a cutover time. We'll verify the details and prep your account for the switch for that time.
4. **Transition/Cutover** — Once SCIM is enabled for your account, it will no longer be possible to add or manage users in the UI. We recommend being ready to test quickly with a few sample users to verify that roles are mapping correctly.
### Role Mapping
Freeplay uses a strict **1:1 mapping** between your directory groups and Freeplay user roles. To assign a role to a user, you must add them to **one and only one** of the specific groups listed below.
If a user is in multiple Freeplay groups at a single time, it is undefined which role from that set they will assume.
#### Default Supported Group Names
If you want a more seamless setup for role management, you can name your IdP groups like the following and we will directly sync the users into Freeplay via a strict string match on these values.
| IdP Group Name | Freeplay Role |
| ------------------ | ------------- |
| `freeplay_admin` | Admin |
| `freeplay_user` | User |
| `freeplay_analyst` | Analyst |
| `freeplay_guest` | Guest |
#### Custom Group Names
If you prefer to define your own group names, the 1:1 mapping still applies. You will need to provide us with a list of your relevant IdP groups and how each one should map to Freeplay's user roles (Admin, User, Analyst, Guest).
This is a manual process today. Changes to your group/group names without informing us may result in users being demoted to the Analyst (default role). Use [support@freeplay.ai](mailto:support@freeplay.ai) for any desired changes here.
For example:
| IdP Group Name | Freeplay Role |
| ----------------- | ------------- |
| `MyCoApp-admin` | Admin |
| `MyCoApp-dev` | User |
| `MyCoApp-analyst` | Analyst |
| `MyCoApp-vendor` | Guest |
Users that are synced to Freeplay with a group name that does not match one of the strings above (or with no group) will default to the **Analyst** role.
### Managing Access and Limitations
Once Directory Sync is enabled for your account, user management behaviors change significantly. Please review the following rules to ensure a smooth workflow.
#### Pre-existing Users
* When first setting up the directory to sync with Freeplay, if you already have users in the Freeplay system they will be retained. As noted below, they will no longer be editable via the Freeplay app.
* **Existing users will retain their assigned role.** In order to ensure your IdP is in sync with what's in Freeplay, you should manually add those existing users to the appropriate group in your IdP.
* **To get a list of the users with their roles:** You can either get them from the UI or from the API. A call to `GET /api/v2//users` will return all active users in the account. Note that to call this endpoint the API key must be associated with a user with the Account Admin role.
* Their role will change if they're added to other groups, or updated.
#### Directory as Source of Truth
* **No UI Creation** — You will no longer be able to create new users inside the Freeplay app UI. All users must be provisioned via SCIM.
* **Locked Profile Data** — Users' First Name, Last Name, and Email are set during provisioning and cannot be edited within Freeplay.
* **Permissions** — User roles are locked to their directory group. You cannot change a user's role manually in the Freeplay UI.
#### Changing User Roles
Freeplay determines a user's role based on **the most recent event received from your directory** (i.e. "last event wins"). We do not support mapping multiple groups to a single user.
**To change a user's role (e.g., from User to Admin):**
Role changes are no longer allowed inside the app. This is to prevent out-of-sync issues. Roles are now handled like so:
1. **Remove** the user from the old group (`freeplay_user`) **first**
2. **Add** the user to the new group (`freeplay_admin`) **second**
If you add a user to a new group before removing them from the old one, the "remove" event for the old group may arrive last. This will leave the user with no valid group, causing them to revert to the default **Analyst** role. This can be easily adjusted by deactivating/reactivating that user's group setting.
#### Project Membership
SCIM manages **Roles** (system-level permissions), but it does not manage Freeplay **Project** membership.
* **Private Projects** — Access to private projects must still be granted manually within the Freeplay UI by a Project admin.
* **Guest Users** — Since the Guest role usually implies restricted access (e.g., for contractors), you must explicitly add them to specific Projects in the Freeplay UI for them to see any content.
### Technical Limitations
* **Timing** — We process events on a 5-minute interval. This means that role changes can take up to 5 minutes to process. Please be aware that role changes will not be instantaneous.
* **Single Group Only** — Users should belong to only one Freeplay-mapped group at a time. If a user is added to multiple groups, the role will reflect the most recent "Add" event received. This can also be fixed on your side by removing and re-adding the user to the highest privileged group that's desired. (e.g. if a user is added to IdP-User and IdP-Admin, but the user event arrived later, they'll be set as IdP-User. Simply re-add them to IdP-Admin to re-grant the admin permissions here).
* **No `memberOf` Support** — We do not support the use of `memberOf` SAML assertions to update permissions on login. Permissions are updated only via SCIM directory events.
# Token Cost Estimates
Source: https://docs.freeplay.ai/account-setup/token-cost-estimates
Understand how Freeplay calculates and displays token cost estimates to help you track and manage LLM API spend.
## Overview
Cost estimates in Freeplay give you a clear view of how you are spending your API tokens. Freeplay breaks down costs by project, prompt template, and evaluation so you can trace your spend and identify where costs are highest.
Freeplay also provides rough estimates of evaluation costs before you run them, helping you make informed decisions about your testing and evaluation strategy.
## Where to find cost estimates
You can view cost information in two places within Freeplay.
### Account settings --> Usage
The [**Usage** tab](https://app.freeplay.ai/settings/members) in Account Settings is the primary source of truth for costs and cost estimates. Navigate to **Account Settings** and select the **Usage** tab.
This page provides two key metrics:
* **App spend** - estimated spend based on logged token counts to Freeplay. This reflects the cost of running your application's LLM calls, not money spent in or through Freeplay itself.
* **Spend via Freeplay** - the total spend from activity within the Freeplay platform. This includes:
* Online evaluations and auto-categories
* Test runs
* Prompt playground usage
* [AI features](/core-concepts/ai/ai-features)
### Estimated evaluation costs
On the evaluations page, each evaluation displays an estimated cost. This estimate gives you a rough idea of how much it costs to run that evaluation over your completions.
The estimate is calculated as follows:
1. Takes the average input and output tokens for the specific prompt template or agent (including input variables) from the last week
2. Calculates the average cost based on those token counts
This is a rough estimate and is subject to change based on volume, sampling rate, and other factors.
## How costs are calculated
Freeplay bases all cost estimates on token usage. Token costs are calculated using up-to-date pricing information from each provider.
Cost estimates may vary if you:
* Use different LLM providers
* Have custom deployment configurations (e.g., Azure OpenAI, AWS Bedrock)
* Have negotiated pricing agreements with providers
### Custom per-token costs
For certain providers such as LiteLLM, you can provide custom input and output per-token cost information. This allows Freeplay to reflect your actual costs more accurately when you have special pricing or use self-hosted models.
To configure custom token costs, see [LiteLLM Proxy](/account-setup/configure-litellm-proxy-models-in-freeplay).
## Frequently asked questions
* **Model selection** -- switching to a smaller or less expensive model for tasks that do not require the most capable model
* **Prompt optimization** -- reducing token counts by writing more concise prompts
* **Sampling rate** -- adjusting the sampling rate for online evaluations to run them on a subset of completions rather than all of them
* **Evaluation frequency** -- running test evaluations less frequently or on smaller datasets during development
Freeplay uses publicly available pricing from each LLM provider to estimate the cost per input and output token. These rates are updated regularly to reflect the latest published pricing.
If you use self-hosted models, custom deployments, or have negotiated pricing with providers, the default per-token costs may not reflect your actual spend. For providers that support it (such as LiteLLM), you can configure custom per-token costs in Freeplay to get more accurate estimates. For other providers, treat the estimates as a relative benchmark for comparing costs across prompts and models.
# Bulk Create Agent Test Cases
Source: https://docs.freeplay.ai/api-reference/agent-datasets/bulk-create-agent-test-cases
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/bulk
Add multiple test cases to an agent dataset in a single request. Use for batch imports from production traces.
Maximum 100 test cases per request.
# Bulk Delete Agent Test Cases
Source: https://docs.freeplay.ai/api-reference/agent-datasets/bulk-delete-agent-test-cases
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/bulk
Remove multiple agent test cases in a single request. Use for batch cleanup operations.
Maximum 100 test cases per request.
# Create Agent-Level Dataset
Source: https://docs.freeplay.ai/api-reference/agent-datasets/create-agent-level-dataset
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/agent-datasets
Create a new dataset for agent-level testing. Use to organize test cases for evaluating full agent workflows.
# Delete Agent Dataset
Source: https://docs.freeplay.ai/api-reference/agent-datasets/delete-agent-dataset
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agent-datasets/{dataset_id}
Archive an agent dataset and its test cases. Use when retiring datasets no longer needed.
# Delete Agent Test Case
Source: https://docs.freeplay.ai/api-reference/agent-datasets/delete-agent-test-case
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/{test_case_id}
Remove a test case from an agent dataset. Use to clean up invalid or outdated test cases.
# Get Agent Dataset
Source: https://docs.freeplay.ai/api-reference/agent-datasets/get-agent-dataset
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets/{dataset_id}
Retrieve an agent dataset's metadata by ID. Use to check dataset configuration or compatible agents.
# Get Agent Test Case
Source: https://docs.freeplay.ai/api-reference/agent-datasets/get-agent-test-case
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/{test_case_id}
Retrieve a specific agent test case by ID. Use to inspect test case details or debug test run results.
# List Agent Datasets
Source: https://docs.freeplay.ai/api-reference/agent-datasets/list-agent-datasets
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets
Retrieve all agent-level datasets in a project. Use to discover available datasets for agent test runs.
`page_size` defaults to 30, maximum 100.
# List Agent Test Cases
Source: https://docs.freeplay.ai/api-reference/agent-datasets/list-agent-test-cases
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases
Retrieve all test cases (row-level values) in an agent dataset. Use to review dataset contents or export for analysis.
`page_size` defaults to 30, maximum 100.
# Update Agent Dataset
Source: https://docs.freeplay.ai/api-reference/agent-datasets/update-agent-dataset
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/agent-datasets/{dataset_id}
Modify an agent dataset's name, description, or compatible agents. Use to evolve datasets as agent requirements change.
# Update Agent Test Case
Source: https://docs.freeplay.ai/api-reference/agent-datasets/update-agent-test-case
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/agent-datasets/{dataset_id}/test-cases/{test_case_id}
Modify an existing agent test case's inputs, outputs, or metadata. Use to correct errors or update expected outputs.
# Create Agent
Source: https://docs.freeplay.ai/api-reference/agents/create-agent
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/agents
Create a new agent in your project. Agents are used to organize and group related sessions and traces.
# Delete Agent
Source: https://docs.freeplay.ai/api-reference/agents/delete-agent
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/agents/{agent_id}
Delete an agent from your project. This operation cannot be undone.
# Get Agent
Source: https://docs.freeplay.ai/api-reference/agents/get-agent
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agents/{agent_id}
Retrieve details for a specific agent by ID.
# List Agents
Source: https://docs.freeplay.ai/api-reference/agents/list-agents
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/agents
Retrieve a paginated list of agents for your project. Optionally filter by name.
`page_size` defaults to 30, maximum 100.
# Update Agent
Source: https://docs.freeplay.ai/api-reference/agents/update-agent
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/agents/{agent_id}
Update the name of an existing agent.
# Add Project Member
Source: https://docs.freeplay.ai/api-reference/configuration/add-project-member
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/members
Grant a user access to the project with a specified role. Requires project admin role.
# Create Environment
Source: https://docs.freeplay.ai/api-reference/configuration/create-environment
https://app.freeplay.ai/openapi.json post /api/v2/environments
Create a new deployment environment. Use to add environments beyond Freeplay defaults, like staging or feature branches.
# Create Project
Source: https://docs.freeplay.ai/api-reference/configuration/create-project
https://app.freeplay.ai/openapi.json post /api/v2/projects
Create a new project in your workspace. Use to organize prompts, datasets, and observability data by team or application.
# Create User
Source: https://docs.freeplay.ai/api-reference/configuration/create-user
https://app.freeplay.ai/openapi.json post /api/v2/users
Create a new user in your workspace. Requires account admin role. Use for automated user provisioning or SCIM integrations.
# Delete Environment
Source: https://docs.freeplay.ai/api-reference/configuration/delete-environment
https://app.freeplay.ai/openapi.json delete /api/v2/environments/{environment_id}
Remove an environment. Use when retiring deployment targets no longer in use.
# Delete Project
Source: https://docs.freeplay.ai/api-reference/configuration/delete-project
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}
Permanently delete a project and all its associated data. Requires project admin role.
# Delete User
Source: https://docs.freeplay.ai/api-reference/configuration/delete-user
https://app.freeplay.ai/openapi.json delete /api/v2/users/{user_id}
Remove a user from your workspace. Requires account admin role. Use for offboarding or access revocation.
# Get Project
Source: https://docs.freeplay.ai/api-reference/configuration/get-project
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}
Retrieve the current project's details. Use to check project settings like spend limits or data retention.
# Get User
Source: https://docs.freeplay.ai/api-reference/configuration/get-user
https://app.freeplay.ai/openapi.json get /api/v2/users/{user_id}
Retrieve a user's details by ID. Requires account admin role. Use to check user settings or role assignments.
# List Environments
Source: https://docs.freeplay.ai/api-reference/configuration/list-environments
https://app.freeplay.ai/openapi.json get /api/v2/environments
Retrieve all deployment environments in your workspace. Use to discover available environments for prompt deployment.
`page_size` defaults to 30, maximum 100.
# List Project Members
Source: https://docs.freeplay.ai/api-reference/configuration/list-project-members
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/members
Retrieve all users with access to the project and their roles. Use for access management or auditing.
`page_size` defaults to 30, maximum 100.
# List Projects
Source: https://docs.freeplay.ai/api-reference/configuration/list-projects
https://app.freeplay.ai/openapi.json get /api/v2/projects/all
Retrieve all projects accessible to the current user. Use to discover available projects or build project selection UIs.
# List Users
Source: https://docs.freeplay.ai/api-reference/configuration/list-users
https://app.freeplay.ai/openapi.json get /api/v2/users
Retrieve all users in your workspace. Requires account admin role. Use for user management or access auditing.
Set include_deleted=true to include soft-deleted users.
# Remove Project Member
Source: https://docs.freeplay.ai/api-reference/configuration/remove-project-member
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/members/{user_id}
Revoke a user's access to the project. Requires project admin role.
# Update Environment
Source: https://docs.freeplay.ai/api-reference/configuration/update-environment
https://app.freeplay.ai/openapi.json patch /api/v2/environments/{environment_id}
Rename an existing environment. Use when consolidating or reorganizing deployment targets.
# Update Project
Source: https://docs.freeplay.ai/api-reference/configuration/update-project
https://app.freeplay.ai/openapi.json put /api/v2/projects/{project_id}
Modify project settings like name, visibility, or resource limits. Requires project admin role.
# Update Project Member
Source: https://docs.freeplay.ai/api-reference/configuration/update-project-member
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/members/{user_id}
Change a user's role in the project. Requires project admin role.
# Update User
Source: https://docs.freeplay.ai/api-reference/configuration/update-user
https://app.freeplay.ai/openapi.json patch /api/v2/users/{user_id}
Modify a user's name, role, or profile settings. Requires account admin role.
# Create Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-code-evaluation-criteria
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/code
Create a new code evaluation criteria with an initial auto-deployed version.
# Create Code Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-code-evaluation-criteria-version
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions
Create a new version of an existing code evaluation criteria. The version is not deployed until published.
# Create Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-evaluation-criteria
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria
Create a new evaluation criteria in a project. Use to set up criteria for evaluating LLM outputs in test runs or online evaluations.
# Create Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/create-evaluation-criteria-version
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions
Create a new version of an evaluation criteria. Use to iterate on evaluation criteria configuration while preserving the version history.
# Delete Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/delete-code-evaluation-criteria
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}
Permanently delete a code evaluation criteria and all its versions.
# Delete Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/delete-evaluation-criteria
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}
Permanently delete an evaluation criteria and all its versions. Use when retiring criteria no longer needed.
# Delete Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/delete-evaluation-criteria-version
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions/{version_id}
Permanently delete an evaluation criteria version. Use to clean up draft or unused versions.
# Deploy Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/deploy-evaluation-criteria-version
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions/{version_id}/deploy
Activate an evaluation criteria version for use in evaluations. Use to promote a tested version to production.
# Disable Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/disable-code-evaluation-criteria
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/disable
Deactivate a code evaluation criteria without deleting it.
# Disable Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/disable-criteria
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/disable
Deactivate an evaluation criteria without deleting it. Use to temporarily pause the use of a given criteria in evaluations.
# Enable Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/enable-code-evaluation-criteria
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/enable
Activate a code evaluation criteria for use in evaluations.
# Enable Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/enable-criteria
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/enable
Activate an evaluation criteria for use in test runs and online evaluations.
# Execute Evals for Completion
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/execute-evals-for-completion
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/completions/{completion_id}/evaluations
Queue evaluations to run against a specific completion. Evaluations are executed asynchronously.
Use the `eval_types` field to control which evaluations run:
- `all` (default): Run both LLM-as-judge and code evaluations
- `llm-as-judge`: Run only LLM-based evaluations
- `code`: Run only code-based evaluations
Optionally pass `criteria_ids` to run specific evaluation criteria. When provided, those
criteria are executed regardless of whether they have already been evaluated. When omitted,
only criteria that have not yet been evaluated for this completion will run.
# Execute Evals for Trace
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/execute-evals-for-trace
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/traces/{trace_id}/evaluations
Queue evaluations to run against a specific trace. Evaluations are executed asynchronously.
Use the `eval_types` field to control which evaluations run:
- `all` (default): Run both LLM-as-judge and code evaluations
- `llm-as-judge`: Run only LLM-based evaluations
- `code`: Run only code-based evaluations
Optionally pass `criteria_ids` to run specific evaluation criteria. When provided, those
criteria are executed regardless of whether they have already been evaluated. When omitted,
only criteria that have not yet been evaluated for this trace will run.
# Get Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-code-evaluation-criteria
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}
Retrieve a code evaluation criteria and its latest version, including eval code.
# Get Code Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-code-evaluation-criteria-version
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions/{version_id}
Retrieve a specific version of a code evaluation criteria.
# Get Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-evaluation-criteria
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}
Retrieve an evaluation criteria's configuration by ID. Use to inspect criteria settings or get the latest version ID.
# Get Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/get-evaluation-criteria-version
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions/{version_id}
Retrieve a specific evaluation criteria version. Use to inspect version configuration or compare with other versions.
# List Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-code-evaluation-criteria
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code
Retrieve all code evaluation criteria in a project. Optionally filter by target type and ID.
`page_size` defaults to 30, maximum 100.
# List Code Evaluation Criteria Versions
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-code-evaluation-criteria-versions
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions
Retrieve all versions of a code evaluation criteria. `page_size` defaults to 30, maximum 100.
# List Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-evaluation-criteria
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria
Retrieve all evaluation criteria in a project. Use to discover available criteria or build criteria management UIs.
`page_size` defaults to 30, maximum 100.
# List Evaluation Criteria Versions
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/list-evaluation-criteria-versions
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/versions
Retrieve all versions of an evaluation criteria. Use to view version history or compare changes over time.
`page_size` defaults to 30, maximum 100.
# Publish Code Evaluation Criteria Version
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/publish-code-evaluation-criteria-version
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}/versions/{version_id}/publish
Deploy a version, undeploying the currently deployed version.
# Reorder Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/reorder-criteria
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/id/{criteria_id}/reorder
Change the display order of evaluation criteria in the Freeplay UI. Use to organize criteria in a logical sequence for human review workflows.
# Update Code Evaluation Criteria
Source: https://docs.freeplay.ai/api-reference/evaluation-criteria/update-code-evaluation-criteria
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/evaluation-criteria/code/{criteria_id}
Update criteria-level metadata. Does not create a new version.
# Get Insights
Source: https://docs.freeplay.ai/api-reference/insights/get-insights
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/insights
Retrieve a paginated list of insights for a project. Use to discover insights or filter by prompt template or agent.
`page_size` defaults to 30, maximum 100.
# Create Model
Source: https://docs.freeplay.ai/api-reference/models/create-model
https://app.freeplay.ai/openapi.json post /api/v2/models
Create or upsert a custom model configuration for the account.
# Delete Model
Source: https://docs.freeplay.ai/api-reference/models/delete-model
https://app.freeplay.ai/openapi.json delete /api/v2/models/{model_id}
Delete a custom model configuration from the account.
# Get Model
Source: https://docs.freeplay.ai/api-reference/models/get-model
https://app.freeplay.ai/openapi.json get /api/v2/models/{model_id}
Retrieve details for a specific custom model by ID.
# List Models
Source: https://docs.freeplay.ai/api-reference/models/list-models
https://app.freeplay.ai/openapi.json get /api/v2/models
Retrieve a paginated list of custom models configured for the account.
`page_size` defaults to 30, maximum 100.
# Update Model
Source: https://docs.freeplay.ai/api-reference/models/update-model
https://app.freeplay.ai/openapi.json put /api/v2/models/{model_id}
Update an existing custom model configuration.
# Add Completion Feedback
Source: https://docs.freeplay.ai/api-reference/observability/add-completion-feedback
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completion-feedback/id/{completion_id}
Record end-user feedback on a completion. Use to capture thumbs up/down ratings or custom feedback attributes. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/customer-feedback)
The `freeplay_feedback` field must be "positive" or "negative". Additional custom fields are supported.
# Add Trace Feedback
Source: https://docs.freeplay.ai/api-reference/observability/add-trace-feedback
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/trace-feedback/id/{trace_id}
Record end-user feedback on a trace. Use to capture feedback on conversation turns or agent workflow outcomes. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/customer-feedback)
The `freeplay_feedback` field must be "positive" or "negative". Additional custom fields are supported.
# Delete Session
Source: https://docs.freeplay.ai/api-reference/observability/delete-session
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/sessions/{session_id}
Permanently delete a session and all associated completions. Use when removing test data or honoring data deletion requests.
# Record Completion
Source: https://docs.freeplay.ai/api-reference/observability/record-completion
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/completions
Log an LLM completion with its prompt, response, and metadata. This is the primary endpoint for observability. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/recording-completions)
Sessions are created implicitly—just generate a session_id client-side (UUID v4). Optionally provide your own completion_id too to avoid waiting for Freeplay's response.
# Record Trace
Source: https://docs.freeplay.ai/api-reference/observability/record-trace
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/sessions/{session_id}/traces/id/{trace_id}
Create or update a trace within a session. Use to group related completions in agent workflows. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/traces)
Generate the trace_id client-side (UUID v4) to avoid round-trip latency.
# Update Completion
Source: https://docs.freeplay.ai/api-reference/observability/update-completion
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completions/{completion_id}
Append messages or evaluation results to an existing completion. Use for streaming responses or adding post-hoc evaluation metrics.
# Update Session Metadata
Source: https://docs.freeplay.ai/api-reference/observability/update-session-metadata
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/id/{session_id}/metadata
Merge custom metadata into an existing session. Use to enrich sessions with post-hoc context like user ID or business metrics. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/sessions)
# Update Trace by ID
Source: https://docs.freeplay.ai/api-reference/observability/update-trace-by-id
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/{session_id}/traces/id/{trace_id}
Update a trace's output, metadata, feedback, eval results, and/or test run info by its trace ID.
# Update Trace by OTEL Span ID
Source: https://docs.freeplay.ai/api-reference/observability/update-trace-by-otel-span-id
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/{session_id}/traces/otel-span-id/{otel_span_id_hex}
Update a trace's output, metadata, feedback, eval results, and/or test run info by its OpenTelemetry span ID (hex string).
Note: OTEL spans are mapped to Freeplay Traces, so we use the OTEL span ID to identify Traces.
# Update Trace Metadata
Source: https://docs.freeplay.ai/api-reference/observability/update-trace-metadata
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/sessions/{session_id}/traces/id/{trace_id}/metadata
Merge custom metadata into an existing trace. Use to enrich traces with post-hoc context.
# Bulk Create Prompt Test Cases
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/bulk-create-prompt-test-cases
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/bulk
Add multiple test cases to a dataset in a single request. Use for batch imports, e.g. from CSV or production logs.
Maximum 100 test cases per request.
# Bulk Delete Prompt Test Cases
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/bulk-delete-prompt-test-cases
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/bulk
Remove multiple test cases in a single request. Use for batch cleanup operations.
Maximum 100 test cases per request.
# Create Prompt-Level Dataset
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/create-prompt-level-dataset
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-datasets
Create a new dataset for prompt-level testing. Use to organize test cases for evaluating individual prompts.
# Delete Prompt Dataset
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/delete-prompt-dataset
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}
Archive a prompt dataset and its test cases. Use when retiring datasets no longer needed.
# Delete Prompt Test Case
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/delete-prompt-test-case
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/{test_case_id}
Remove a test case from a dataset. Use to clean up invalid or outdated test cases.
# Get Prompt Dataset
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-dataset
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/id/{dataset_id}
# Get Prompt Dataset
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-dataset-1
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}
Retrieve a prompt dataset's metadata by ID. Use to check dataset configuration or input schema.
# Get Prompt Dataset by Name
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-dataset-by-name
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/name/{dataset_name}
# Get Prompt Test Case
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/get-prompt-test-case
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/{test_case_id}
Retrieve a specific test case (dataset row) by ID. Use to inspect test case details or debug test run results.
# List Prompt Datasets
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-datasets
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets
Retrieve all prompt-level datasets in a project. Use to discover available datasets for test runs.
`page_size` defaults to 30, maximum 100.
# List Prompt Test Cases
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-test-cases
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/id/{dataset_id}/test-cases
`page_size` defaults to 10, maximum 10.
# List Prompt Test Cases
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-test-cases-1
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases
Retrieve all row-level test cases in a prompt dataset. Use to review dataset contents or export for analysis.
`page_size` defaults to 30, maximum 100.
# List Prompt Test Cases by Dataset Name
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/list-prompt-test-cases-by-dataset-name
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/datasets/name/{dataset_name}/test-cases
`page_size` defaults to 10, maximum 10.
# Update Prompt Dataset
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/update-prompt-dataset
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}
Modify a prompt dataset's name, description, or input schema. Use to evolve datasets as prompt requirements change.
# Update Prompt Test Case
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/update-prompt-test-case
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/prompt-datasets/{dataset_id}/test-cases/{test_case_id}
Modify an existing test case's inputs, output, or metadata. Use to change ground truth output values or correct errors.
# Upload Prompt Test Cases
Source: https://docs.freeplay.ai/api-reference/prompt-datasets/upload-prompt-test-cases
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/datasets/id/{dataset_id}/test-cases
Add test cases to a dataset using the legacy upload format. Use for bulk imports with input/output pairs.
Maximum 100 examples per request.
# Cancel a pending or in-progress prompt optimization job.
Source: https://docs.freeplay.ai/api-reference/prompt-optimization/cancel-a-pending-or-in-progress-prompt-optimization-job
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-optimization-jobs/{job_id}/cancel
Only jobs that have not yet completed can be cancelled.
# Get prompt optimization job status and details.
Source: https://docs.freeplay.ai/api-reference/prompt-optimization/get-prompt-optimization-job-status-and-details
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-optimization-jobs/{job_id}
Returns the current status of the prompt optimization job, including progress information
for polling. When complete, includes the optimized_version_id and optionally
test_run_id and comparison_id if run_test_after_optimization was True.
# List prompt optimization jobs for the project.
Source: https://docs.freeplay.ai/api-reference/prompt-optimization/list-prompt-optimization-jobs-for-the-project
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-optimization-jobs
Returns a paginated list of prompt optimization jobs, sorted by creation date (newest first).
Optionally filter by status.
`page_size` defaults to 30, maximum 100.
# Start a new prompt optimization job.
Source: https://docs.freeplay.ai/api-reference/prompt-optimization/start-a-new-prompt-optimization-job
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-optimization-jobs
Creates an asynchronous job that optimizes the specified prompt template version.
The job will analyze examples from the dataset and generate an improved prompt.
When `run_test_after_optimization` is True (default), the job will also run
baseline and optimized test runs and create a comparison.
Poll GET /prompt-optimization-jobs/{job_id} to check job status.
# Create prompt template
Source: https://docs.freeplay.ai/api-reference/prompt-templates/create-prompt-template
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates
Create a new prompt template without any versions. Use when you need to reserve a template name before adding versions.
# Create prompt template version by ID
Source: https://docs.freeplay.ai/api-reference/prompt-templates/create-prompt-template-version-by-id
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions
Create a new version of an existing prompt template. Freeplay assigns a random ID. Use when you have the template ID from a previous API call.
**Version Creation Semantics:**
A new version is created only if there is no existing version with identical:
- Content (prompt messages)
- Model
- LLM parameters
- Version name (if provided)
- Version description
- Tool schema
- Output schema
If an identical version exists in any of the target environments, that version is reused and its environments are updated to match the requested environments. This ensures you never create duplicate versions with the same configuration.
**Environment Deployment:**
- When creating a new version, it will be deployed to the specified environments (or "latest" if none specified)
- When reuploading an existing prompt template with a new environment, the environments on that prompt template version will be updated
These checks are performed so that you can safely upload prompt templates without worrying that you will be creating duplicate versions. This is especially useful in development workflows where your prompts are stored in code.
# Create prompt template version by name
Source: https://docs.freeplay.ai/api-reference/prompt-templates/create-prompt-template-version-by-name
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/name/{template_name}/versions
Create a new version of a prompt template, referenced by a string or name you define. [See SDK method.](https://docs.freeplay.ai/freeplay-sdk/prompts)
Set `create_template_if_not_exists=true` to auto-create the template if it doesn't exist.
**Version Creation Semantics:**
A new version is created only if there is no existing version with identical:
- Content (prompt messages)
- Model
- LLM parameters
- Version name (if provided)
- Version description
- Tool schema
- Output schema
If an identical version exists in any of the target environments, that version is reused and its environments are updated to match the requested environments. This ensures you never create duplicate versions with the same configuration.
**Environment Deployment:**
- When creating a new version, it will be deployed to the specified environments (or "latest" if none specified)
- When reuploading an existing prompt template with a new environment, the environments on that prompt template version will be updated
These checks are performed so that you can safely upload prompt templates without worrying that you will be creating duplicate versions. This is especially useful in development workflows where your prompts are stored in code.
# Delete prompt template
Source: https://docs.freeplay.ai/api-reference/prompt-templates/delete-prompt-template
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-templates/id/{template_id}
Archive a prompt template and all its versions. Use when retiring templates that are no longer needed. This is a soft delete.
# Delete prompt template version
Source: https://docs.freeplay.ai/api-reference/prompt-templates/delete-prompt-template-version
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id}
# Get all prompt templates by environment
Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-all-prompt-templates-by-environment
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/all/{environment}
# Get bound prompt template by name and environment
Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-bound-prompt-template-by-name-and-environment
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/name/{name}
# Get bound prompt template version by ID
Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-bound-prompt-template-version-by-id
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id}
# Get prompt template
Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-prompt-template
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/id/{template_id}
Retrieve a prompt template's metadata by ID. Use to check if a template exists or get its latest version ID.
# Get prompt template by name and environment
Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-prompt-template-by-name-and-environment
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/name/{name}
# Get prompt template version by ID
Source: https://docs.freeplay.ai/api-reference/prompt-templates/get-prompt-template-version-by-id
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id}
# List prompt template versions
Source: https://docs.freeplay.ai/api-reference/prompt-templates/list-prompt-template-versions
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions
Retrieve all versions of a prompt template. Use to view version history or compare changes over time.
`page_size` defaults to 30, maximum 100.
# List prompt templates
Source: https://docs.freeplay.ai/api-reference/prompt-templates/list-prompt-templates
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/prompt-templates
Retrieve all prompt templates in a project. Use to discover available templates or build template management UIs.
`page_size` defaults to 30, maximum 100.
# Update environment for prompt template version
Source: https://docs.freeplay.ai/api-reference/prompt-templates/update-environment-for-prompt-template-version
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/prompt-templates/id/{template_id}/versions/{version_id}/environments
Deploy a prompt version to one or more environments. Use to promote versions through dev, staging, and production.
# Update prompt template
Source: https://docs.freeplay.ai/api-reference/prompt-templates/update-prompt-template
https://app.freeplay.ai/openapi.json patch /api/v2/projects/{project_id}/prompt-templates/id/{template_id}
Rename a prompt template. Use when refactoring template names across your codebase.
# List Review Queues
Source: https://docs.freeplay.ai/api-reference/review-queues/list-review-queues
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/review-queues
Retrieve a paginated list of review queues for the project, including
creator, assignees, and review progress counts.
`page_size` defaults to 30, maximum 100.
# Create Saved Search
Source: https://docs.freeplay.ai/api-reference/saved-searches/create-saved-search
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/session-filters
Create or upsert a saved search for the project.
# Delete Saved Search
Source: https://docs.freeplay.ai/api-reference/saved-searches/delete-saved-search
https://app.freeplay.ai/openapi.json delete /api/v2/projects/{project_id}/session-filters/{filter_id}
Delete a saved search from the project.
# Get Saved Search
Source: https://docs.freeplay.ai/api-reference/saved-searches/get-saved-search
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/session-filters/{filter_id}
Retrieve details for a specific saved search by ID.
# List Saved Searches
Source: https://docs.freeplay.ai/api-reference/saved-searches/list-saved-searches
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/session-filters
Retrieve a paginated list of saved searches for the project.
`page_size` defaults to 30, maximum 100.
# Update Saved Search
Source: https://docs.freeplay.ai/api-reference/saved-searches/update-saved-search
https://app.freeplay.ai/openapi.json put /api/v2/projects/{project_id}/session-filters/{filter_id}
Update an existing saved search.
# Get All Completion Statistics
Source: https://docs.freeplay.ai/api-reference/search-&-analytics/get-all-completion-statistics
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completions/statistics
Retrieve aggregate evaluation statistics across all prompts for a date range. Use for dashboard metrics or trend analysis.
Maximum date range is 30 days.
# Get Completion Statistics
Source: https://docs.freeplay.ai/api-reference/search-&-analytics/get-completion-statistics
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/completions/statistics/{prompt_template_id}
Retrieve evaluation statistics for a specific prompt template. Use to track quality metrics for individual prompts.
Maximum date range is 30 days.
# List Sessions
Source: https://docs.freeplay.ai/api-reference/search-&-analytics/list-sessions
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/sessions
Retrieve sessions with their completions, ordered by most recent first. Use to display conversation history or analyze LLM usage patterns. Traces are referenced by ID but not expanded.
Filter by custom metadata using query parameters prefixed with `custom_metadata.` (e.g., `custom_metadata.user_id=123`).
`page_size` defaults to 10, maximum 100.
Prefer using the `/search/sessions` endpoint for more advanced filtering.
# Search Completions
Source: https://docs.freeplay.ai/api-reference/search-&-analytics/search-completions
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/search/completions
Query LLM completions using advanced filters. Use to find specific
prompts and responses, filter by evaluation results or metadata, prompt
templates, latency, and more.
Supports pagination and optional inclusion of all child traces and
completions within the session, using the `include_children` parameter.
`page_size` defaults to 30, maximum 100.
For filter operators and examples, see: https://docs.freeplay.ai/openapi/search-api-operators
# Search Sessions
Source: https://docs.freeplay.ai/api-reference/search-&-analytics/search-sessions
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/search/sessions
Query sessions using advanced filters. Use to find specific
conversations or filter by metadata.
Supports pagination and optional inclusion of all traces and completions
within the session, using the `include_children` parameter.
`page_size` defaults to 30, maximum 100.
For filter operators and examples, see: https://docs.freeplay.ai/openapi/search-api-operators
# Search Traces
Source: https://docs.freeplay.ai/api-reference/search-&-analytics/search-traces
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/search/traces
Query traces using advanced filters. Use to find specific trace
executions, filter by metadata, or analyze trace patterns.
Supports pagination and optional inclusion of all child traces and
completions within the session, using the `include_children` parameter.
`page_size` defaults to 30, maximum 100.
For filter operators and examples, see: https://docs.freeplay.ai/openapi/search-api-operators
# Create Tag
Source: https://docs.freeplay.ai/api-reference/tags/create-tag
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/tags
Create a new tag in the project. If color is not provided, it will be auto-assigned.
# List Project Tags
Source: https://docs.freeplay.ai/api-reference/tags/list-project-tags
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/tags
Retrieve all tags in a project. Optionally filter by entity type (e.g., 'dataset').
# Create Test Run
Source: https://docs.freeplay.ai/api-reference/test-runs/create-test-run
https://app.freeplay.ai/openapi.json post /api/v2/projects/{project_id}/test-runs
# Get Test Run Results
Source: https://docs.freeplay.ai/api-reference/test-runs/get-test-run-results
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/test-runs/id/{test_run_id}
# List Test Runs
Source: https://docs.freeplay.ai/api-reference/test-runs/list-test-runs
https://app.freeplay.ai/openapi.json get /api/v2/projects/{project_id}/test-runs
`page_size` defaults to 100, maximum 100.
# AI Features
Source: https://docs.freeplay.ai/core-concepts/ai/ai-features
How Freeplay uses AI to accelerate your product improvement workflow.
Surface patterns and root causes from your evaluation data, human reviews, and test runs
Score individual completions and traces using LLMs to evaluate your AI outputs at scale
Create better evaluation criteria with AI-powered suggestions and prompt drafts for LLM judges
Classify logs to reveal usage patterns and understand how users interact with your AI
Get AI-generated suggestions for improved prompts based on your production data
## Overview
All AI features in Freeplay work by calling LLM APIs to analyze your data. They are designed to work with different models and to use your API keys and model preferences, based on your account settings.
## Managing AI feature settings
### Disabling specific features
Individual AI features can be controlled through their respective configuration:
* **Model-graded evaluations**: Disable per evaluation by turning off or setting sample rate to zero
* **Eval Creation Assistant**: This is an on-demand feature that only runs when creating evals
* **Auto-categorization**: Disable per auto-category by turning off or setting sample rate to zero
* **Prompt optimization**: This is an on-demand feature that only runs when triggered
* **Review Insights**: Runs automatically during review; disable via the Insights toggle in Project Settings > AI Features
* **Evaluation Insights**: Runs weekly; disable via the Insights toggle in Project Settings > AI Features
### Cost considerations
AI features consume tokens from the selected LLM provider. Costs depend on:
* Which features you use and how frequently
* The volume of data being analyzed
* The models being used (more capable models typically cost more)
When Freeplay Keys are enabled, Freeplay covers the cost of AI feature usage. When using your own API keys, costs are billed directly to your provider account.
Token usage for AI features is tracked separately from your application's LLM usage and is visible in the Usage dashboard. If you're using your own API keys, monitor this usage and consider adjusting feature sampling rates if costs are higher than expected.
## Related pages
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure LLM judges for automated scoring
* [Auto-Categorization](/core-concepts/evaluations/auto-categorization) — Set up automatic content classification
* [Review Queues](/core-concepts/review-queues) — Organize human review workflows and surface insights
* [Model Management](/account-setup/model-management) — Configure LLM providers and API keys
# Overview
Source: https://docs.freeplay.ai/core-concepts/ai/ai-insights
How Freeplay's AI Insights agent analyzes your data to surface actionable findings and improvement opportunities.
AI Insights can be toggled off in **Project Settings > AI Features**.
Freeplay's AI Insights is an intelligence layer that sits at each transition point in your agent development workflow, helping you move from raw data to decisions. A background AI agent analyzes your labeled data — evaluation results, human annotations, and test runs — to generate findings. These findings help you understand the **why** behind performance changes and what to do about it.
The goal is simple: every time you start spending time in Freeplay, you should quickly be able to spot the next most impactful thing you can do to improve your AI system. Then after every labeling session or experiment you run, you should quickly know what to do next.
## The problem Insights solve
You can quickly log hundreds of thousands or millions of traces. You can run LLM judges or other metrics to score those logs, and visualize those in dashboards showing pass/fail rates. Those might tell you that you have a problem, but scores rarely tell you how to fix anything.
Freeplay helps you track evaluation performance over time but raw metrics alone stop short of helping you understand **why** performance changed or **what to do about it**. When you try to decide what to actually improve, you're stuck asking the same questions: Why is that metric failing? What should I do to fix it? Where should I start first?
AI Insights closes that gap by proactively generating findings that add a layer of interpretation on top of your raw metrics — turning scores into direction. See our [blog post](https://freeplay.ai/blog/automated-insights-for-ai-agents) for more information.
## How Insights work
Insights come from Freeplay's own AI agent that analyzes your data from multiple sources and generates actionable findings.
### Where Insights run
Insights run in two places across the Freeplay platform, each representing a decision point in the AI quality workflow:
| Location | Helps you answer |
| :------------------ | :------------------------------------------------------------------------------------------------------------ |
| **Production logs** | Where should I focus? What's broken that I didn't know about? |
| **Human reviews** | What patterns are emerging across my team's annotations? What are the root causes of issues people have seen? |
Each of these represents a moment where you need to interpret lots of data and decide what to do next — exactly the kind of work AI is good at.
A single completion or trace can be tagged with more than one insight.
### Types of Insights
Freeplay generates two types of Insights, each tied to a different data source and decision point in your workflow:
* [**Evaluation Insights -**](/core-concepts/ai/evaluation-insights) Analyze production logs scored by LLM-as-a-judge evaluations. Run on a weekly cadence to surface systemic issues across your logged data.
* [**Review Insights -**](/core-concepts/ai/review-insights) Analyze human annotations in real time. Every note, label, or evaluation triggers the agent to identify patterns and group them into themes.
### What data insights use
Each type of insight is based on different types of information within Freeplay. Here are the sources for each insight:
| | [**Evaluation**](/core-concepts/ai/evaluation-insights) Insights | Review Insights |
| ------------------------ | ---------------------------------------------------------------- | --------------- |
| Model-graded evaluations | :check | :check |
| Human evaluations | X | :check |
| Human notes and lables | X | :check |
| Logged data | :check | :check |
AI Insights does not currently use code evaluations or auto-categorizations as input sources.
## Viewing and using Insights
AI Insights are viewable from the **Home page** or the **Insights** tab within your project.
### Refining Insights
You can fine-tune insights to improve their accuracy and usefulness:
* **Update the name** — providing a more descriptive name can slightly adjust and refine the grouping of tagged records
* **Edit the description** — adding more details, specific errors, or patterns found in your analysis helps fine-tune what the insight captures
### Resolving Insights
The goal of insights is to **resolve them and surface new ones**. Insights can help lead your team towards solving the key issues in your product. Once an issue is identified and fixed, the insight will start to lose traction as no new information is added to it.
## Insights and the data flywheel
Insights provide a clean path towards understanding the **why** behind errors. Some outcomes of insights include:
* **Prompt improvements** — actionable suggestions for how to modify your prompts
* **New evaluations** — generating new LLM-as-a-judge evals based on discovered patterns
* **Deeper investigation** — surfacing issues that people might be missing
Combined with [Review Queues](/core-concepts/review-queues), [prompt optimization](/core-concepts/ai/prompt-optimization), and automated evaluations, Insights help your data flywheel operate smoothly.
## Related resources
* [AI Features Overview](/core-concepts/ai/ai-features) — Overview of all AI-powered features in Freeplay
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure the LLM judges that feed into Insights
* [Review Queues](/core-concepts/review-queues) — Set up human review workflows that generate Review Insights
* [Automated Insights for AI Agents](https://freeplay.ai/blog/automated-insights-for-ai-agents) — Blog post with more detail on the vision behind Insights
# Auto-Categorization
Source: https://docs.freeplay.ai/core-concepts/ai/auto-categorization
Use AI to automatically tag and classify your production logs based on categories you define.
Auto-categorization uses AI to automatically tag and classify your production logs based on categories you define. This adds a layer of intelligence that helps you understand usage patterns and identify trends.
## How it works
1. You define category types that align with your business needs (e.g., product areas, user intent types, issue categories)
2. For each category, you provide a clear name and description
3. As logs flow through Freeplay, the AI classifies them according to your categories
4. Categories appear in the observability dashboard for filtering and analysis
## Use cases
* **Usage analysis**: Understand what types of questions users ask most frequently
* **Issue identification**: Track which product areas generate the most problems
* **Dataset curation**: Filter logs by category to build targeted test datasets
* **Review queue creation**: Focus review efforts on specific categories
## Configuration
Auto-categorization is configured at the prompt template or agent level, similar to other evaluations:
1. Navigate to your prompt template or agent
2. Create a new evaluation with type **Multi-select**
3. Enable auto-categorization and define your categories
4. Each category needs a name (max 32 characters) and description (max 500 characters)
5. Configure whether items can be tagged with multiple categories or just one
**Best practice:** Auto-categorization works best with clear, mutually exclusive categories. If you see many items tagged as "Other" or miscategorized, refine your category descriptions.
[Learn more about auto-categorization →](/core-concepts/evaluations/auto-categorization)
# Eval Creation Assistant
Source: https://docs.freeplay.ai/core-concepts/ai/eval-creation-assistant
Create better evaluation criteria with AI-powered suggestions and prompt drafts for LLM judges.
Writing effective evaluation prompts can be challenging, especially for teams new to LLM-based quality assessment. Freeplay's Eval Creation Assistant uses AI to help you draft better evals faster—whether you're starting from scratch or adapting a template.
## How it works
The Eval Creation Assistant helps in two ways:
**Create custom evals from scratch**: Start with the basic question you want to answer about your AI's output. The assistant will:
1. Help you refine your evaluation question to be clear and measurable
2. Suggest improvements to your eval structure
3. Automatically draft a model-graded eval prompt tailored to your specific prompts and data
**Adapt from templates**: Choose from common evaluation templates like Answer Faithfulness (for RAG), Similarity, Toxicity, or Tone. The assistant will:
1. Automatically customize the template to match your prompt structure
2. Reference the correct input variables from your prompts
3. Generate a ready-to-use eval prompt with one click
Because Freeplay knows your prompt structure and has access to real-world examples from your logs, the assistant can generate eval prompts that are specific to your context rather than generic templates.
## Use cases
* **Getting started quickly**: Teams new to evals can create their first evaluations without prior experience
* **Adopting best practices**: Start with industry-standard eval patterns and customize them for your needs
* **Cross-functional collaboration**: Product managers, analysts, and domain experts can contribute to eval creation without writing code
## Using the assistant
1. Navigate to your prompt template or agent
2. Go to the **Evaluations** section
3. Choose **Create your own** or select from the template library
4. For custom evals: Enter your evaluation question and follow the AI's suggestions
5. For templates: Select a template and the AI will automatically adapt it to your prompt
6. Test the generated eval against sample data
7. Use the [alignment flow](/practical-guides/creating-and-aligning-model-graded-evals) to validate that the eval matches human judgment
Even when using templates, the AI adapts them to your specific prompt variables and data structure—so you get truly customized evals, not just generic prompts.
**Best practice:** If you're new to writing evals or unsure where to start, use the Eval Creation Assistant's "Create your own" option. Describe what you want to evaluate in plain language, and the AI will generate a custom eval prompt tailored to your specific prompts and use case.
# Evaluation Insights
Source: https://docs.freeplay.ai/core-concepts/ai/evaluation-insights
AI-powered analysis of your production evaluation data to surface issues and improvement opportunities.
Evaluation Insights analyze your production log data to surface issues you might not catch from dashboards alone.
## How they work
Freeplay proactively analyzes logged data that has [model-graded evaluations](/core-concepts/evaluations/model-graded-evaluations) applied. The agent reviews these logs and identifies key patterns across the data. Here is the general flow:
1. Freeplay collects evaluation results over a time period (requiring at least 10 logs with evaluation data)
2. The AI analyzes the logs, looking for patterns in:
* Poor-scoring outputs and their common characteristics
* Correlation between different evaluation criteria
* Input patterns that tend to produce poor results
3. The agent then reviews these results, assigns, creates or updates existing insights to properly assign and group the data
For each insight, you get a clear description of the problem, an easy link to the underlying traces that back it up, and the number of matching records as a proxy for scale and impact to help you prioritize what matters most.
## When they run
Evaluation Insights run on a **weekly cadence**, generating findings every Monday morning based on the previous week's data.
Evaluation Insights can be disabled in **Project Settings > AI Features**.
## Related resources
* [AI Insights Overview](/core-concepts/ai/ai-insights) — How Freeplay's AI Insights work across the platform
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure the LLM judges that feed into Insights
# Model-Graded Evaluations
Source: https://docs.freeplay.ai/core-concepts/ai/model-graded-evaluations
Use AI to automatically score your AI outputs based on evaluation criteria you define.
Model-graded evaluations (also called LLM judges or auto-evaluations) use AI to automatically score your AI outputs based on criteria you define. This is the foundation of automated quality assessment in Freeplay.
## How it works
When you configure a model-graded evaluation:
1. You define the evaluation criteria with a name, question, and scoring type (Yes/No, 1-5 scale, etc.)
2. You write instructions explaining what the LLM should evaluate and provide a rubric with scoring guidelines
3. Freeplay generates a structured prompt that includes your criteria, the completion being evaluated, and relevant context
The LLM then scores each completion according to your rubric and provides an explanation for its decision.
## Use cases
* **Production monitoring**: Automatically sample and evaluate a subset of production traffic
* **Batch testing**: Run evaluations across entire datasets during test runs
* **Quality gates**: Identify outputs that fail specific quality thresholds
## Configuration
Model-graded evaluations are configured at the prompt template or agent level:
1. Select the **Evaluations** tab from the menu and then select **New Evaluation**
2. Set the target to your **prompt/agent** to evaluate and the type to **Model-graded**
3. Create your own or select from a pre-configured example
4. Write instructions that reference your prompt variables (e.g., `{{inputs.context}}`, `{{output}}`)
5. Define a rubric that maps scores to specific behaviors
Use Freeplay's alignment tools to compare auto-evaluation scores against human labels and iteratively improve your evaluation prompts.
**Best practice:** Model-graded evaluations are the foundation for many other AI features. Prompt optimization and Evaluation Insights both work better when you have well-configured evaluations generating data. Start here before enabling other AI features.
[Learn more about model-graded evaluations →](/core-concepts/evaluations/model-graded-evaluations)
# Prompt Optimization
Source: https://docs.freeplay.ai/core-concepts/ai/prompt-optimization
Use AI to analyze your production data, evaluation results, and customer feedback to suggest improved prompts.
Prompt optimization uses AI to analyze your production data, evaluation results, and customer feedback to suggest improved prompts. It can also help update prompts when switching between models.
## How it works
1. You select a prompt template version to optimize and choose a dataset or set of evaluated sessions
2. You configure what data sources to use:
* **Human labels**: Scores and feedback from your team's reviews
* **Customer feedback**: Direct feedback captured from end users
* **Best practices**: Provider-specific prompting guides (OpenAI or Anthropic)
3. You can optionally provide specific instructions about what to improve
4. Freeplay's AI analyzes the data and generates:
* An optimized prompt template
* An explanation of changes made
* A description of the new version
## Use cases
* **Prompt iteration**: Get AI-suggested improvements based on where your current prompt is failing
* **Model migration**: Update prompts optimized for one model to work well with another
* **Data-driven improvement**: Use production signals to guide prompt changes
## Configuration
Prompt optimization is available from the prompt template editor:
1. Open a prompt template and select a version
2. Click **Optimize** to open the optimization panel
3. Select your data source (dataset or evaluated sessions)
4. Choose which signals to include (labels, feedback, best practices)
5. Optionally add specific instructions
6. Run the optimization
After optimization completes, Freeplay creates a new prompt version and automatically runs a comparative test so you can evaluate the results side-by-side.
Prompt optimization works best with at least 10-20 evaluated examples that include a mix of good and poor outputs.
# Review Insights
Source: https://docs.freeplay.ai/core-concepts/ai/review-insights
Review Insights work alongside your human reviewers to perform real-time root cause analysis.
As your team reviews completions and traces, the insights agent surfaces patterns and groups them into actionable findings.
## How they work
When humans label data in Freeplay the Insights agent analyzes each reviewed item in the background. It identifies common patterns, groups related items into **themes**, and suggests actions based on what it finds.
The inputs to Review Insights include:
* **Human labels** — annotations, notes, and scores from human reviewers
* **LLM-as-a-judge evaluations** — scores and reasoning from your auto-evaluators applied during review
* **Logs** — the completions or traces that were evaluated
## When they run
Review Insights run **anytime a human label is added** to data. Every annotation — notes, human evals, or LLM-as-a-judge evals — triggers the agent to analyze and update insights. When combined with [Review Queues](/core-concepts/review-queues) these review insights can point to key issues in your system.
Review Insights can be disabled in **Project Settings > AI Features**.
Review Insights themes are generated automatically and may occasionally be too broad or too narrow. Regularly review themes and use merge/prune actions to keep them useful.
## Related resources
* [AI Insights Overview](/core-concepts/ai/ai-insights) — How Freeplay's AI Insights work across the platform
* [Review Queues](/core-concepts/review-queues) — Set up human review workflows that generate Review Insights
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) — Configure the LLM judges that feed into Insights
# Curating Useful Datasets for Testing & Evaluation
Source: https://docs.freeplay.ai/core-concepts/datasets/dataset-curation
Learn strategies for building high-quality datasets that accurately represent real-world usage.
## What are datasets and why are they important?
Testing and evaluation are key aspects of the LLM development cycle. Your test quality and reliability are a function of two primary components: your Evaluators and your Datasets. In this guide, we are going to focus on Dataset curation.
Simply put, datasets are collections of inputs and outputs that you can use to test your LLM systems. Each example's **output** represents either a golden response (the ideal answer) or a captured failure case from production. Including outputs is strongly recommended — they are what evaluations compare new LLM responses against during test runs. For more details, see [Understanding the Output Field](/core-concepts/datasets/datasets#understanding-the-output-field).
Datasets are important because for your tests to be truly informative your datasets need to accurately represent the issues and situations your LLM systems face in the wild.
We’ll cover how we at Freeplay think strategically about building datasets and then look tactically at how to curate datasets inside of Freeplay.
## Dataset Curation Strategy
Once your team has decided what product or feature you want to build, a common next step is curating your datasets. Many teams will build a single dataset of various scenarios for an LLM feature and instinctively stop there. While one dataset is a fantastic starting point, teams quickly realize that multiple datasets are crucial for effective testing and experimentation. Broadly speaking there are two types of datasets: Targeted datasets and Broad-based datasets.
### Targeted datasets
Targeted datasets are datasets that are focused on a narrowly defined issue or situation.
For example, let’s say you’re working on an e-commerce use case in which we are using an LLM to answer customer questions about their orders.
The pipeline has two components:
1. First, the LLM generates a SQL query from the user question
2. Then, the LLM uses the results of that query to generate an answer
To test this pipeline, you might create a targeted dataset called “Query Hallucinations”. This dataset would be a collection of examples in which the model hallucinates a table name. You might create another targeted dataset called “Delivered Orders”, which collects examples where the user asked about an order that was already delivered. The first dataset is focused on a technical failure point and the second dataset is focused on a specific customer situation, orders that have already been delivered. Both datasets are narrowly focused.
When iterating on LLMs it’s often useful to focus on one specific problem at a time. Targeted datasets allow you to quickly iterate over your key areas of concern.
### Broad-based datasets
Broad-based datasets are datasets that include a wide array of examples and are not focused on any specific issue or situation.
The classic example of a broad-based dataset would be the “golden set.” A golden set is a dataset consisting of examples hand-curated by humans to be the ideal output for some given input.
Unlike a targeted dataset, broad-based datasets contain a variety of situations meant to capture the totality of cases your LLM system needs to be able to handle. This kind of dataset is often used for benchmarking or regression testing.
These two types of datasets can then be used together during experimentation. Continuing our e-commerce example from earlier, let’s say you’re focused on reducing hallucinations in SQL query generation. You can first focus on making prompt and model changes for that specific issue, frequently testing against your targeted dataset along the way. Then once you think you have a fix in hand, you test the new config against your broad-based dataset to ensure you haven’t regressed on other dimensions.
## Anatomy of a Freeplay Dataset
Now that you’re familiar with the primary types of datasets, let’s take a look at what a Freeplay dataset consists of.
* Name - ex. “Query Hallucinations”
* Description - ex. “Sessions where the model referenced an invalid table”
* Prompt Compatibility - A set of inputs that your dataset will be compatible with.
* Examples - Examples are combinations of inputs and outputs that you as the user save to a dataset. The output should be either a golden response or a captured failure case — see [Understanding the Output Field](/core-concepts/datasets/datasets#understanding-the-output-field) for guidance.
## Curating Datasets in Freeplay
### Step 1: Create a new Dataset
Navigate to the Datasets tab and select “Create dataset”. Give your dataset a Name and Description, then set your Prompt Compatibility. You’ll need to decide what prompt(s) you want your dataset to be compatible with. Compatibility is determined by the prompt’s input variables. Datasets can be compatible with multiple prompts as long as those prompts share at least one common input variable.
Alternatively, you can create a new Dataset directly from a completion or trace by clicking “+ Dataset” and then hitting the + icon. Note, if you create the dataset this way, prompt compatibility will be intuited for you based on the completion you create the dataset from.
### Step 2: Add Examples
There are a number of ways to add examples to a dataset.
**Add from completion**
From any completion you can hit “Add to dataset” to create a new example from that completion. This is often a big part of the human review flow, as reviewers are labeling data they can actively be building datasets as well.
#### Curating Samples from Production Data
When adding new samples directly from Observability, you have the option to manually curate them to be representative samples in your dataset. When you open up the "+ Dataset" modal, it will give you the ability to curate any of the inputs, variables, history and output by adding additional messages, tool calls, multimedia and more. This allows you to add quality samples to your dataset making it even more useful for testing different parts of the product such as failure modes, successful cases and more!
**Bulk add from observability**
You can bulk add completions to a dataset by going to the Observability tab, toggling to the completions view in the table and then selecting the completions you want to add. Often we will see users filter on things like eval values, customer feedback, or other metrics and bulk adding completions from there.
**Upload examples**
If you have existing examples you can upload them to Freeplay via JSONL. Navigate to the Datasets tab, select your dataset, and click upload.
You can read more about formatting the JSONL file [here](/core-concepts/datasets/datasets#uploading-data-using-jsonl).
**Manually write examples**
You can also write examples directly in the UI. From any dataset click “Create an example” and a form will appear where you can write a new example by hand. You’ll enter values for each input variable as well as the output. It’s okay to leave any of these blank if it makes sense for your example.
### Step 3: Run a Test against your Dataset
After you’ve created a dataset you can run a batch test with any of your compatible prompts. Batch tests can be kicked off either from the Freeplay app or via the SDK.
To run a batch test from the UI go to the Tests tab and click “Run Test”. From there you can configure the test by selecting the prompt version you want to test and the dataset you want to test with.
To run a batch test from the SDK see the docs [here](/freeplay-sdk/test-runs).
### Step 4: Managing Datasets
You can manage your dataset on an ongoing basis in the Datasets tab. Here you can add, edit, and delete examples
### Bonus: Use your Dataset in the Playground
When editing a prompt in the playground you can pull in examples from your dataset and run them in real time to test your changes.
In the prompt editor click the folder icon and load in examples
## Key Takeaways
Dataset curation is an often overlooked part of the LLM development cycle. Your testing is only as good as your underlying datasets. Having a rich collection of datasets empowers developers to iterate faster and ultimately deliver better, higher-quality AI features for your customers. Freeplay helps facilitate that dataset curations process in a fully integrated platform.
# Datasets
Source: https://docs.freeplay.ai/core-concepts/datasets/datasets
Build and manage test datasets to power evaluations, test runs, and fine-tuning workflows.
Datasets in Freeplay are an essential part of organizing data to test your LLM systems. They can also be used to curate data for human review or fine-tuning. Datasets are the foundation of [Test Runs](/core-concepts/test-runs/test-runs) in Freeplay.
**API Reference**: Freeplay supports two types of datasets:
* **Prompt Datasets** (for component-level testing)
* **Agent Datasets** (for end-to-end testing)
Each is automatically built to enforce schemas that maintain compatibility with your separate prompts and agents.
A key benefit of using Freeplay to curate Datasets is that it's seamless to save new examples that you observe in real-world testing or production to existing Datasets. This keeps the data fresh and representative of the actual use of your application.
Datasets can be created to test LLM systems across a variety of scenarios, such as:
* **Golden Set:** For detecting regressions vs. your ideal ground truth
* **Failure Cases:** For tracking failures you observe and testing in the future to confirm they are fixed
* **Red Teaming:** For managing adversarial test cases and confirming appropriate behavior by your system
* **Random Samples:** For representative testing across a distributed set of values
Instructions on how to save observed data or upload data are below.
## Understanding the output field
Every dataset entry has an **output** field. While not strictly required, we strongly recommend including an output for each example — it plays a central role in evaluations and test runs, and examples without an output have limited utility for testing.
There are two primary ways to use the output field:
* **Golden output:** The output represents the ideal, correct response for the given inputs. This is common in golden sets and broad-based datasets where you want to benchmark new prompt versions against a curated standard. When used in test runs or in the playground, these outputs can be viewed to see how the newly generated data compares to the ideal output.
* **Failure case:** The output captures a real failure observed in production — such as a hallucination, incorrect answer, or off-tone response. This is useful for building targeted datasets that track known issues so you can confirm they are fixed in future prompt versions.
The output can come from any source — uploaded files, completions saved from observed logs, or manually written examples. When saving completions from production logs, you have the option to edit the output before saving, which allows you to curate it into a golden response or preserve it as a failure case depending on your testing goals.
# Curating Datasets
Datasets in Freeplay can be curated in one of two ways: by saving completions that are recorded to Freeplay straight from the Sessions view, or by uploading existing test cases to a Dataset.
## Saving Data from Recorded Sessions
While working with recorded Sessions or Traces in Freeplay, if you encounter values that are relevant for future testing, you can save it directly. You will be given the option to curate the inputs and outputs before saving to the dataset. This can be useful if you want to make this sample represent a specific type of data sample such as a golden or failure case. This can be done at the trace or completion view.
To do this, simply:
* Click `+ Dataset` above the completion/trace view
* Optionally, make adjustments to the inputs, history or outputs
* Select the relevant dataset(s)
* Optionally, click the `+` button to create a new dataset from this menu
### Bulk Add
You can also select multiple completions or traces at once and add a large group of completions to a dataset at one time, even across pages.
* Select the "Completions" or "Traces" view on Observability (instead of Sessions)
* Click the radio buttons in the table for the rows you want
## Adding Metadata to Dataset Entries
Metadata can now be added to entries in your datasets, allowing you to store additional information with each entry.
To add or edit metadata for a dataset entry:
1. Navigate to a specific dataset entry
2. Click the "Edit" option in the dropdown menu
3. In edit mode, you'll see a dedicated "Metadata" section at the top of the entry
4. Add customizable key-value pairs such as:
* Customer identifiers (e.g., "customerId": "2382721")
5. Click "Add Metadata" to create additional fields as needed
6. Click "Save" to store your changes
# Uploading Datasets
If you have existing data that is relevant to use for testing prompts in Freeplay, you can upload it directly as a JSONL or CSV file. Both formats support the same fields.
### Prompt Template Datasets
| Field | Required | Description |
| ------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `inputs.` | **Yes** (at least one) | Each input variable must be prefixed with `inputs.` (e.g., `inputs.question`, `inputs.job_info`). These map directly to `{{variable_name}}` in your prompt template. |
| `history` | No | A JSON array of previous messages representing the conversation history. See [Tool Calls in History](#tool-calls-in-history) below for supported message formats. |
| `output` | Yes (empty ok) | The expected or ideal output for the given inputs — either a golden response or a captured failure case. See [Understanding the Output Field](#understanding-the-output-field). |
| `metadata.` | No | Additional information on each data sample. Metadata columns must start with `metadata.` (e.g., `metadata.employee_id`, `metadata.source`). |
For more details on variable usage, see our [Advanced Prompt Templating](/core-concepts/prompt-management/advanced-prompt-templating-using-mustache) guide.
**Tool Calls in History:** The `history` field supports tool call and tool result messages, allowing you to upload conversation histories that include function calling interactions. Assistant messages can contain `tool_call` content blocks, and user messages can contain `tool_result` content blocks. See the samples below for examples.
## How to Upload
1. **Navigate to the Dataset** and click the **Upload** button.
2. **For CSV uploads**, click **Download CSV Template** in the bottom-left corner of the upload dialog to get a file with the correct column names for your dataset.
3. **Format your data** according to the JSONL or CSV format below, ensuring that each entry aligns with your selected prompt template.
4. **Upload your file.** If there are any formatting issues, Freeplay will flag them and block the upload with a warning.
## JSONL Format
Each line in the file must be a valid JSON object flattened to a single line. Be sure to append the filename with `.jsonl`. Note that JSONL is NOT normal JSON — see [jsonlines.org](https://jsonlines.org/) for the specification.
```json jsonl theme={null}
{"inputs": {"job_info": {"title": "Senior Software Engineer", "level": "IC4", "department": "Engineering", "hire_date": "2021-03-15", "manager": "David Park"}}, "history": [{"role": "user", "content": "Can you pull up the compensation summary for Sarah Chen?"}, {"role": "assistant", "content": [{"type": "tool_call", "id": "call_001", "name": "lookup_employee", "arguments": {"employee_name": "Sarah Chen"}}]}, {"role": "user", "content": [{"type": "tool_result", "tool_call_id": "call_001", "content": "{\"employee_id\": \"EMP001\", \"name\": \"Sarah Chen\", \"current_salary\": 178000, \"currency\": \"USD\"}"}]}, {"role": "assistant", "content": "I found Sarah Chen's profile. Let me generate her compensation brief now."}], "output": "Summary\n\nSarah Chen — Senior Software Engineer (IC4), Engineering. Base: $178,000 USD.", "metadata": {"employee_id": "EMP001", "name": "Sarah Chen", "department": "Engineering"}}
{"inputs": {"job_info": {"title": "Marketing Manager", "level": "M1", "department": "Marketing", "hire_date": "2021-11-01", "manager": "Rachel Adams"}}, "history": [{"role": "user", "content": "I need to prep for Emma's comp review."}, {"role": "assistant", "content": [{"type": "tool_call", "id": "call_002", "name": "lookup_employee", "arguments": {"employee_name": "Emma Clarke"}}, {"type": "tool_call", "id": "call_003", "name": "get_market_data", "arguments": {"market": "London Tech", "role": "Marketing Manager"}}]}, {"role": "user", "content": [{"type": "tool_result", "tool_call_id": "call_002", "content": "{\"employee_id\": \"EMP007\", \"name\": \"Emma Clarke\", \"current_salary\": 83000, \"currency\": \"GBP\"}"}, {"type": "tool_result", "tool_call_id": "call_003", "content": "{\"median_salary\": 80000, \"p75_salary\": 92000, \"currency\": \"GBP\"}"}]}, {"role": "assistant", "content": "I've pulled Emma's profile and London market benchmarks. Generating her brief now."}], "output": "Summary\n\nEmma Clarke — Marketing Manager (M1), Marketing. Base: £83,000 GBP. Market median: £80,000.", "metadata": {"employee_id": "EMP007", "name": "Emma Clarke", "department": "Marketing"}}
```
## CSV Format
Your CSV must use specific column header prefixes for Freeplay to correctly parse your data. Any columns that don't follow these conventions will be ignored.
```csv theme={null}
inputs.job_info,history,output,metadata.employee_id,metadata.name,metadata.department
"{""title"": ""Senior Software Engineer"", ""level"": ""IC4"", ""department"": ""Engineering"", ""hire_date"": ""2021-03-15"", ""manager"": ""David Park""}","[{""role"": ""user"", ""content"": ""Can you pull up the compensation summary for Sarah Chen?""}, {""role"": ""assistant"", ""content"": [{""type"": ""tool_call"", ""id"": ""call_001"", ""name"": ""lookup_employee"", ""arguments"": {""employee_name"": ""Sarah Chen""}}]}, {""role"": ""user"", ""content"": [{""type"": ""tool_result"", ""tool_call_id"": ""call_001"", ""content"": ""{\""employee_id\"": \""EMP001\"", \""name\"": \""Sarah Chen\"", \""current_salary\"": 178000, \""currency\"": \""USD\""}""}]}, {""role"": ""assistant"", ""content"": ""I found Sarah Chen's profile. Let me generate her compensation brief now.""}]","Summary
Sarah Chen — Senior Software Engineer (IC4), Engineering. Base: $178,000 USD.","EMP001","Sarah Chen","Engineering"
"{""title"": ""Marketing Manager"", ""level"": ""M1"", ""department"": ""Marketing"", ""hire_date"": ""2021-11-01"", ""manager"": ""Rachel Adams""}","[{""role"": ""user"", ""content"": ""I need to prep for Emma's comp review.""}, {""role"": ""assistant"", ""content"": [{""type"": ""tool_call"", ""id"": ""call_002"", ""name"": ""lookup_employee"", ""arguments"": {""employee_name"": ""Emma Clarke""}}, {""type"": ""tool_call"", ""id"": ""call_003"", ""name"": ""get_market_data"", ""arguments"": {""market"": ""London Tech"", ""role"": ""Marketing Manager""}}]}, {""role"": ""user"", ""content"": [{""type"": ""tool_result"", ""tool_call_id"": ""call_002"", ""content"": ""{\""employee_id\"": \""EMP007\"", \""name\"": \""Emma Clarke\"", \""current_salary\"": 83000, \""currency\"": \""GBP\""}""}, {""type"": ""tool_result"", ""tool_call_id"": ""call_003"", ""content"": ""{\""median_salary\"": 80000, \""p75_salary\"": 92000, \""currency\"": \""GBP\""}""}]}, {""role"": ""assistant"", ""content"": ""I've pulled Emma's profile and London market benchmarks. Generating her brief now.""}]","Summary
Emma Clarke — Marketing Manager (M1), Marketing. Base: £83,000 GBP. Market median: £80,000.","EMP007","Emma Clarke","Marketing"
,,,,,"
```
# Dataset Compatibility
We've found that it's important to allow for relatively flexible compatibility rules to accommodate complex prompting strategies. The following compatibility rules may be important to know:
* Compatibility for testing is based on the input `{{variable_names}}` in your prompt templates. These must match with the key names in your Datasets.
* A Dataset is treated as compatible if **one or more** key names match for a given prompt template. This is important so that datasets can be treated as compatible even when some variable names are optional in practice. (See [Advanced Prompt Templating Using Mustache](/core-concepts/prompt-management/advanced-prompt-templating-using-mustache))
* Datasets can be used across multiple prompt templates in a Project, as long as at least one variable name is shared. For instance, if you have four prompt templates that all use the variable `{{question}}`, then any Dataset that contains values for `{{question}}` will be compatible.
***
What's Next
* [Curating Useful Datasets for Testing & Evaluation](/core-concepts/datasets/dataset-curation)
* [Test Runs](/core-concepts/test-runs/test-runs)
# Auto-categorization
Source: https://docs.freeplay.ai/core-concepts/evaluations/auto-categorization
Automatically tag and classify your production logs to understand usage patterns and identify trends.
# Overview
Auto-categorization provides teams with an additional layer of intelligence about their AI systems by automatically tagging incoming logs with specified categories. This feature adds valuable context to your production data, helping product and engineering teams understand how their AI applications are being used and where improvements are needed.
## Why Use Auto-Categorization
For teams building AI products, understanding usage patterns and identifying trends is essential. Auto-categorization reveals what types of questions users ask, which product areas generate the most activity, and where your system might need attention.
When combined with evaluations, auto-categorization helps pinpoint exactly which types of inputs challenge your system. Tracking categories over time reveals usage trends, helps identify emerging patterns, and provides product teams with actionable insights about feature adoption and user behavior.
## Implementing Auto-categorization
Auto-categorization works at both the agent and completion level. Start by defining category types that align with your business needs—such as product areas (API/SDK, Observability, Prompt Management) or user intent types (Technical Support, Billing, Product Information).
### Creating Effective Categories
To set up a new category, follow these steps:
1. **Name your categorization scheme** - Choose a descriptive name for the overall categorization (e.g., "Product Area")
2. **Select relevant references** - Choose which variables (input, output, history) the LLM should consider when categorizing
3. **Define individual categories** - Add specific categories with clear, distinct descriptions
4. **Configure multi-category options** - Decide whether items can be tagged with multiple categories or just one
## Best Practices
Keep descriptions clear and distinct - Category names are limited to 32 characters and descriptions to 500 characters. Make each category clearly distinguishable:
* **Category:** "API/SDK"
* **Description:** "Questions about API endpoints, authentication, SDK installation, code integration, or programmatic access to the platform"
Use variable references strategically - While you can reference variables like input or output in your descriptions, these won't be directly injected. Instead, they guide the LLM's attention to relevant parts of the interaction.
## Using Categories in Practice
Once configured, categories appear in the evaluation panel of completions and traces. The ✨ icon reveals the LLM's reasoning for each categorization. Teams can approve correct categorizations for human validation or manually override incorrect classifications.
### Observability & Monitoring
The Observability dashboard transforms your categories into actionable intelligence. Monitor category distributions to understand usage patterns, track feature adoption, or identify areas needing attention. The stacked bar chart visualization makes it easy to see category breakdowns over time.
### Creating Targeted Datasets and Review Queues
Auto-categorization streamlines the process of creating focused subsets for testing and review. By building targeted datasets, teams can create curated and focused datasets that can be used to test specific types of inputs. For review queue creation, this allows product and engineering teams to collaborate and focus on specific parts of the product for review. This helps the team concentrate and focus in order to improve quality of the system. To create and curate these review queues and datasets, simply use observability to search by the specific categories of the auto-categorization, select all the completions, and add to a dataset or review queue!
### Implementation Tips
Start with broad categories - Begin with high-level categorizations before creating more granular subcategories. For example:
* **Documentation** - "Questions about finding, understanding, or using product documentation and guides"
* **Account Management** - "Login issues, password resets, user permissions, or team access concerns"
* **Performance Issues** - "Reports of slow response times, timeouts, high latency, or system availability problems"
Collaborate across teams - Product and engineering teams can jointly define categories to ensure they capture both technical and business-relevant insights.
***
[Code Evaluations](/core-concepts/evaluations/code-evaluations)
[Review Queues](/core-concepts/review-queues)
# Code Evaluations
Source: https://docs.freeplay.ai/core-concepts/evaluations/code-evaluations
Programmatically check outputs for specific patterns, formatting, and key information with deterministic evaluation functions.
## Overview
Code evaluations let you programmatically check outputs for specific patterns, formatting, key information, and more — giving you fast, cheap, and fully reproducible results. Unlike [model-graded evaluations](/core-concepts/evaluations/model-graded-evaluations) that use LLMs as judges, code evaluations are deterministic functions that work with both completions and agents.
Freeplay supports two types of code evaluations:
* **Server-side code evaluations** — Evaluation functions that run directly on Freeplay's servers. You write and manage them in the Freeplay UI, and they execute automatically against your production traffic or during test runs.
* **Client-side code evaluations** — Evaluation functions that run in your own codebase, with results logged to Freeplay via the [SDK](/freeplay-sdk#record-client-side-evals). Client-side evaluations give you full control over execution and can run at any scale without constraints.
Both types of results appear alongside your other evaluations in the observability dashboard, providing an additional layer of insight into your data.
## Server-side code evaluations
Server-side code evaluations run directly on Freeplay's servers, enabling you to compare outputs against any combination of inputs, metadata, and conversation history without writing any integration code.
Server-side code evaluations run in two capacities:
1. **Live monitoring of production sessions** — Freeplay automatically samples a subset of your production traffic and runs code evaluations to give you insight into how your systems are behaving.
2. [**Test Runs**](/core-concepts/test-runs/test-runs) — Test runs allow you to proactively run tests against datasets to measure system performance over time and compare changes side by side.
### Creating a server-side code evaluation
To create a new code evaluation:
1. Navigate to the **Evaluations** section and select **Code**
2. **Select a target** — Choose whether this evaluation targets a prompt template or an agent. This determines which data is available to your evaluation function.
3. **Select a language** — Choose between Python and JavaScript for your evaluation code
4. Give your evaluation a name and configure the output type (boolean or float)
5. Write the code for your evaluation, test against dataset samples and deploy.
### Available values
Once created, your evaluation function has access to the following data:
* **Inputs:** Variables passed to the prompt template as inputs or the input to the agent
* **History:** Previous messages in the conversation
* **Metadata:** Metadata recorded with the completion
* **Output:** The LLM's response text
* **Reference Output:** Expected output from dataset (if available)
For agents, you only have access to Input, Output, Metadata, and Reference Output.
### Output types
Code evaluations support two output types:
* **Boolean** — Returns `true` or `false`, useful for pass/fail checks like schema validation or keyword presence
* **Float** — Returns a numeric value, useful for similarity scores, distance metrics, or percentage-based checks
### Available libraries
Freeplay provides a set of core libraries for each language to help with common comparison and validation tasks. You can view the full list of available imports by clicking the **?** icon in the code editor sidebar.
### Testing your evaluation
Similar to testing in the playground, you can load up to 100 test cases to validate your evaluation before deploying it. The test interface displays variables on the left side and provides several ways to run your evaluation:
* **Run a single test case** — Execute against one test case to quickly verify logic
* **Run all** — Execute against your full set of loaded test cases
Results appear in the **Executed test cases** tab, which indicates any errors and provides runtime logs. This gives you a clean debugging experience — you can see which test cases failed and why, allowing you to iterate and refine your evaluation quickly.
### Using reference output
The `reference_output` parameter is a special case. When your evaluation references this field, it becomes a **test run only** evaluation and cannot run in a live monitoring capacity.
Reference output allows you to compare newly generated output against the expected output stored with a test case. This is useful when you need to verify that changes to your system preserve existing behavior. For example, if you are changing the structure of your output by adding a new key but want to ensure all other outputs remain the same, compare the `output` to the `reference_output` to accomplish this.
Using `reference_output` in your evaluation restricts it to test runs only. You cannot conditionally use it — if it appears in your code, the evaluation will not run in an online capacity.
### Common use cases
* **String matching** — Exact match, regex, or fuzzy string checks
* **Schema validation** — Verify JSON structure, required fields, or data types
* **Output formatting** — Check for expected formatting patterns or constraints
* **Tool call verification** — Validate that the correct tools were used with the right parameters
* **Transcript analysis** — Analyze turn count, token usage, or conversation flow
* **Outcome verification** — Confirm specific business logic conditions are met
## Client-side code evaluations
Client-side code evaluations are evaluation functions that you write and run in your own codebase, then log results to Freeplay. These evaluations give you complete flexibility — you can use any libraries, access external services, and run evaluations at any scale.
Client-side code evaluations are useful for criteria requiring logical expressions, such as JSON schema checks or category assertions, or for pairwise comparisons to an expected output via methods like embedding distance or string similarity. Client-side code evaluations can be added to:
* **Individual sessions** — Run evaluations as part of your application logic and record results alongside completions
* **Test runs executed with the SDK or API** — Include comparisons to ground truth data in batch evaluations
Results you log to Freeplay appear in the UI alongside human and model-graded evaluations. See the [SDK documentation](/freeplay-sdk#record-client-side-evals) for implementation details.
### When to use client-side code evaluations
Client-side code evaluations are the recommended approach when you need to:
* Run evaluations against all of your data without sampling constraints
* Use custom libraries or external services not available in Freeplay's managed environment
* Ensure your system can depend on the evaluated result (i.e., the system relies on the response schema)
* Access private data sources or internal APIs during evaluation
## Viewing code eval results
All code evaluation results appear in the evaluations side panel for both agents and completions under the **Evals** section. Server-side evaluations appear as code evals, while client-side evaluations appear as client evals. Similar to other evaluations, you can use them to filter, set up automations, and view graphs to track their results.
## Frequently asked questions
Code evaluations are best suited for deterministic checks where you want to verify a specific pattern, format, or piece of information in the output. Use them when you have clear, objective criteria that can be expressed in code.
* **Live monitoring** — Use code evaluations for deterministic checks where you want to continuously validate that outputs meet specific criteria in production
* **Test runs** — Use code evaluations when you need to compare inputs to outputs or measure how closely results match expected outputs in a dataset.
* **Server-side code evaluations** are ideal when you want Freeplay to manage execution — they run automatically against sampled production traffic and during test runs with no integration code needed. They also let you test new prompt versions against golden output.
* **Client-side code evaluations** are best when you need full control over execution, want to use custom libraries, need to run evaluations against all of your data without sampling constraints, or your code relies on the result of the evaluation.
Code evaluations are included in your Freeplay plan at no incremental charge, with limits varying by tier. Each evaluation run spins up a dedicated cloud function for execution. To manage usage effectively, avoid setting the sampling rate to 100% for high-traffic applications. [Contact us](/resources/support) for details on limits for your plan.
***
What's next
Now review each evaluation type and then move on to test runs once all your evaluations are configured.
* [Human Labeled Evaluations](/core-concepts/evaluations/human-evaluations)
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations)
* [Evaluations](/core-concepts/evaluations/evaluations)
* [Test Runs](/core-concepts/test-runs/test-runs)
# Evaluations
Source: https://docs.freeplay.ai/core-concepts/evaluations/evaluations
Measure and improve your AI outputs with model-graded, code-based, and human evaluations.
# Evaluations Overview
Evaluation in machine learning is the process of determining a model's performance via a metrics-driven analysis.
Freeplay allows you to incorporate evaluations into your product development lifecycle in a way that is focused on your particular product context or domain. By defining appropriate evaluations for your specific use case, you gain insights that are far more valuable than generic industry benchmarks. Read more on our blog [here](https://freeplay.ai/blog/defining-the-right-evaluation-criteria-for-your-llm-project-a-practical-guide).
Freeplay supports four modes of evaluations that each work together:
* **[Human evaluation](/core-concepts/evaluations/human-evaluations)**: aka "data annotation" or "labeling", where your team can easily review and score results
* **[Model-graded evaluation](/core-concepts/evaluations/model-graded-evaluations)**: using LLMs as a judge for nuanced evaluation criteria instead of humans
* **[Code evaluation](/core-concepts/evaluations/code-evaluations)**: deterministic evaluation functions that run server-side on Freeplay's servers or client-side in your own codebase to validate outputs against specific patterns, formats, and conditions
* **[Auto-categorization](/core-concepts/evaluations/auto-categorization)**: automated tagging of your application logs with specified categories
Some criteria may be appropriate only for human evaluation, while others can benefit from humans working together with model-graded auto-evaluators — giving humans the ability to inspect, confirm or correct any auto-eval results and improve on the model-graded results.
# Configuring Evaluation Criteria
For each of your prompts, you can configure one or more relevant human, model-graded, or code evaluation criteria in Freeplay.
Any evaluation criteria configured in Freeplay can be used for human labeling/annotation, and you can optionally enable model-graded auto-evaluations for relevant criteria too. For example, you might want model-graded evals to score the quality of an LLM response, but you only want humans to be able to leave notes on a completion. [Client-side code evaluations](/core-concepts/evaluations/code-evaluations#client-side-code-evaluations) can be logged to Freeplay directly using our [SDKs](/freeplay-sdk#record-client-side-evals).
# Resources
* [Creating a Testing & Evaluation Process](https://freeplay.ai/blog/prompt-engineering-for-product-managers-part-2-testing-evaluation)
* [Defining the Right Evaluation Criteria](https://freeplay.ai/blog/defining-the-right-evaluation-criteria-for-your-llm-project-a-practical-guide)
***
What's Next
Now review each evaluation type and then move onto test runs once all your evaluations are configured!
* [Code Evaluations](/core-concepts/evaluations/code-evaluations)
* [Human Labeled Evaluations](/core-concepts/evaluations/human-evaluations)
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations)
* [Creating and Aligning Model-Graded Evals](/practical-guides/creating-and-aligning-model-graded-evals) - Practical guide for building effective LLM judges
* [Test Runs](/core-concepts/test-runs/test-runs)
# Human Labeled Evaluations
Source: https://docs.freeplay.ai/core-concepts/evaluations/human-evaluations
Capture expert judgment and subjective quality assessments from your team.
## Human Evaluation/Labeling
## Why Human Evaluation Matters
Human evaluation is essential for measuring dimensions of quality that automated evals can't capture—like cases of high sensitivity, required SME knowledge, or nuanced accuracy in specialized domains. While model-graded and code-based evals scale efficiently, human judgment remains the gold standard for subjective quality measures and for validating that your automated evals are actually measuring what matters.
### Common Use Cases
**Spot-checking production quality** - Regularly sample a subset of production completions to ensure your LLM maintains quality standards. This catches issues that automated evals might miss and helps calibrate your team's understanding of "good" vs. "bad" outputs.
**Building ground truth datasets** - Create labeled datasets that become the foundation for model-graded evaluations. Human labels serve as the "answer key" that trains and validates your automated evaluation layer.
**Measuring subjective dimensions** - Evaluate qualities like helpfulness, empathy, tone appropriateness, or creativity—aspects where human judgment is more reliable than algorithmic scoring.
**Debugging edge cases** - When automated evals flag unusual patterns or when users report issues, human review helps you understand what's actually happening and whether it's a real problem.
**Calibrating automated evals** - Compare human labels to model-graded eval scores to measure alignment. This validates whether your automated evals are trustworthy enough to use at scale.
## Getting Started with Human Evaluation
Freeplay makes it easy for your team to label sessions directly in the platform. Team members can label individual sessions or filter groups of sessions that share common criteria (e.g., weekly spot checks of production data, all completions from a specific user segment, or sessions where automated evals flagged potential issues).
**1. Invite your team** Navigate to **Settings > Account > New user** to add team members. Only Admins can invite new users. Consider inviting domain experts, product managers, or customer success team members—the people who understand quality in your specific context.
**2. Browse and search sessions** Use the **Observability** tab to search sessions based on date ranges, eval scores, user feedback, custom metadata, or any other criteria. This helps you focus human review time on the sessions that matter most.
**3. Apply labels** Navigate to individual sessions and apply labels in the Evaluation section of the sidebar. Hover over tooltips to see the evaluation criteria and instructions you configured when creating the eval. Labels you apply here feed directly into your evaluation analytics and can be used to build training datasets.
**4. Review in batches** For efficiency, add completions and traces from any search to a **Review Queue**. This creates a dedicated workspace where team members can systematically work through records, apply labels, leave comments, and track progress—perfect for regular spot-checking workflows.
### Best Practices
* **Start small**: Begin with a manageable sample size (10-20 sessions) to calibrate your team's understanding of the evaluation criteria
* **Create clear rubrics**: Define specific, actionable criteria in your evaluation instructions so different team members label consistently
* **Track inter-rater reliability**: Have multiple people label the same sessions to measure agreement and refine your criteria
* **Use stratified sampling**: When spot-checking production, sample across different user segments, time periods, or use cases to get representative coverage
* **Close the loop**: Share insights from human evaluation with your engineering team to improve prompts, tune automated evals, or identify training needs
***
What's Next
Now review each evaluation type and then move onto test runs once all your evaluations are configured!
* [Code Evaluations](/core-concepts/evaluations/code-evaluations)
* [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations)
* [Evaluations](/core-concepts/evaluations/evaluations)
* [Test Runs](/core-concepts/test-runs/test-runs)
# Model-Graded Evaluations
Source: https://docs.freeplay.ai/core-concepts/evaluations/model-graded-evaluations
Use LLMs to automatically score and evaluate your AI outputs at scale.
## Model-graded Evaluations
Model-graded auto-evaluations on the Freeplay platform are performed in two capacities:
1. Live Monitoring of Production Sessions
1. Freeplay will automatically sample a subset of your production traffic and run auto evaluations on them to give you insight into how your systems are behaving in the wild.
2. Batch testing via Test Runs
1. Test Runs allow you to proactively run batch tests against Datasets to measure your system performance over time and compare changes side by side. These Test Runs can be executed either via the UI or via the SDK.
## Configuring Model-graded Evaluations in Freeplay
You'll start by configuring evaluation criteria on a prompt template. Go to Prompts > Pick the prompt you want > scroll to the Evaluations section at the bottom.
Each evaluation criteria requires the following components:
1. **Name**: Give the criteria an easy-to-recognize name that will show up in the UI for your team.
2. **Question**: Along with the name, define the guiding question that a human evaluator should address for that criteria. This question should be clear and focused, with a goal to make sure each evaluator understands objectively how to answer the question.
3. **Evaluation Type**: Freeplay currently supports 4 types of evaluation criteria:
1. Yes/No boolean (use "Yes" as the positive value)
2. 1-5 Scale (use "5" as the positive value)
3. Text (free text string, useful for leaving comments or other descriptions of issues)
4. Multi-select (tags/enums, useful for data categorization)
4. **Optionally: Enable model-graded auto-evaluation**: Choose whether you want to configure and run an auto-evaluator for the criteria. More details on setting up model-graded auto-evaluators below.
* Enter a prompt in the "Instructions" section. This is where you detail to the LLM what it is evaluating and importantly indicate which parts of the prompt or output you want to target. This is what dictates the values passed to the LLM at execution time.
* You can target components of your prompt with mustache syntax. In this case we will use `{{inputs.supporting_information}}` to target our retrieved context and `{{output}}` to target the LLM generated output which are the two components we need for this eval.
* Next configure a Rubric, this section details the criteria the LLM will use to make its final decision and can be extremely valuable for generating high quality responses
* Next configure a Rubric, this section details the criteria the LLM will use to make its final decision and can be extremely valuable for generating high quality response.
5. **Optionally: Align your Auto-evaluators**
Prompts for auto-evaluators can take some iteration to get right, just like with other prompt engineering. Freeplay provides functionality to align your Auto Evaluators with human feedback from your team as you're creating your eval criteria. Each time you update the prompt, you can re-run it against a sample of examples and compare to your team's choices.
***
What's Next
Now review each evaluation type and then move onto test runs once all your evaluations are configured!
* [Code Evaluations](/core-concepts/evaluations/code-evaluations)
* [Human Labeled Evaluations](/core-concepts/evaluations/human-evaluations)
* [Evaluations](/core-concepts/evaluations/evaluations)
* [Test Runs](/core-concepts/test-runs/test-runs)
# Saved Searches and Automations
Source: https://docs.freeplay.ai/core-concepts/observability/automations
Use powerful custom filters to explore your logs, then create automated workflows to act on your production data.
Freeplay provides a powerful set of tools to help you monitor, analyze, and take action on your log data. These capabilities build on each other:
1. **Custom Filters** let you define complex queries to find specific data in your logs
2. **Saved Searches** let you save important filters as reusable monitors that you can return to regularly
3. **Automations** let you trigger actions automatically when logs match your saved search criteria
## Custom Filters
When working with large volumes of production data, custom filters let you define complex boolean queries to find exactly the data you care about. Filtering is the foundation for all monitoring and automation capabilities in Freeplay.
You can create filters directly in the Observability dashboard by selecting criteria such as:
* Input and output values
* Evaluation results and scores
* Custom metadata (user IDs, session types, versions of your code, etc.)
* Prompt templates/versions
* Common metrics (e.g. cost, latency, token usage)
Combine multiple criteria with AND/OR logic to build precise queries that surface the completions you need to investigate, review, or act upon. After you apply a filter, it's easy to batch select results and add to a dataset or review queue.
## Saved Searches
Once you've defined a filter that surfaces important data, you can also save it as a **saved search** to return to it easily. Saved searches act as monitors that help you keep track of important states in your logs, like negative customer feedback, failed guardrails, or low scores from auto-evaluators.
Saved searches can either be:
* **Private**: Visible only to you, for personal monitoring and investigation workflows
* **Shared**: Visible to your entire team, so everyone can track important metrics and patterns together
Use saved searches when you have queries you run frequently, want to monitor specific conditions over time, or need to share important monitors with your team.
Once you've created a saved search, you can also use it as the foundation for automations.
## Automations
Automations build on saved searches to take action on filtered data automatically. Once configured, automations run in the background -- continuously monitoring your production traffic and executing actions when logs match your saved search criteria. This allows you to build powerful workflows that ensure the right data gets reviewed, tested, and acted upon without manual effort.
Freeplay supports four types of automations:
Rather than manually searching for problematic completions, automations ensure they're surfaced to your team automatically. Configure which review queue to use, assign specific team members as reviewers (completions are automatically distributed), and set your sampling frequency, limits, and strategy (i.e. random or most recent). This will start adding individual completions or traces to the review queue in the background for you.
**Example use case:** Automatically add guardrailed responses to a review queue so your team can evaluate why guardrails were triggered and identify patterns in edge cases.
This is particularly useful for building datasets from completions that have been reviewed, validated, or scored highly in production. Select which agent or prompt template and dataset to target, and logs will be automatically added to this dataset. This ensures your test coverage grows organically as your system encounters new scenarios in production.
**Example use case:** Automatically add reviewed completions with high eval scores to a golden dataset for regression testing and prompt optimization.
This allows you to run specific evals only on relevant subsets of your data -- saving costs and focusing evaluation effort on what matters most.
Choose which evaluation criteria to run and set your sampling frequency and limits. This is especially useful for running detailed or expensive evals only on completions or traces that meet certain conditions.
**Example use case:** Run detailed evals only on completions that already passed basic checks, or run specialized evals for specific use cases to understand nuanced quality metrics.
Stay informed about important patterns or issues in your production data without constantly monitoring the dashboard.
Configure your notification channel in Slack, set a sampling frequency, and time period for when notifications should trigger.
**Example use case:** Get notified when completions fail a critical evaluation metric within a time period, or when guardrails are triggered.
## Creating an Automation
To create an automation, start by defining a saved search in the Observability dashboard. The saved search determines which logs your automation will act on.
Once you've created a saved search:
1. Click the **"Add automation"** button to configure the automation
2. Give your automation a descriptive name that clearly indicates its purpose (e.g., "Add to Guardrail Review" or "Add to Router Golden Dataset")
3. Select the automation type (Review Queue, Dataset, Run Evals, or Notify)
4. Configure the specific options for your automation type:
* For review queues: select the queue and assignees
* For datasets: choose the target prompt or agent and dataset
* For evals: select the target prompt or agent and which evaluations to run
* For Slack notifications: configure the channel and thresholds
5. Set your sampling frequency (hourly, daily, or weekly)
6. Set the limit for maximum completions per sampling period
7. Choose your sampling strategy (random or most recent)
8. Click **"Save"**
Your automation will now run in the background according to the schedule you configured, continuously processing new completions that match your filter criteria.
## Managing Saved Searches and Automations
All saved searches and their associated automations are visible in the Observability dashboard. From this view, you can edit automation configurations to adjust frequency, limits, or filter criteria. You can also delete saved searches or automations you no longer need.
Monitor automation results regularly to ensure they're capturing the data you expect and taking the intended actions. You can adjust configurations as your needs evolve or as you learn more about your data patterns.
## Best Practices
**Start with specific filters:** The more targeted your filter, the more useful your saved search or automation will be. Broad filters may capture too much irrelevant data, while specific filters surface exactly what you need to monitor, review, or test.
**Use descriptive names:** Name saved searches and automations clearly so your entire team understands their purpose at a glance (e.g., "Low Score Completions" for a saved search, or "Auto-Add Low Scores to Review" for an automation).
**Set appropriate limits:** Start with conservative sampling limits and adjust based on the volume of matching completions. You can always increase limits if you're not capturing enough data, but starting too high may overwhelm review queues or datasets.
**Share important monitors:** When you create a saved search that tracks critical metrics or patterns (like guardrail failures or low customer satisfaction), share it with your team so everyone can monitor these conditions together.
**Combine automation types:** Use multiple automations on the same saved search for different purposes. For example, you might both add low-scoring completions to a review queue AND notify your team when they exceed a threshold.
**Monitor results regularly:** Check your saved searches to understand data patterns and verify that automations are capturing the data you expect. Adjust filter criteria or automation configurations as needed based on what you learn.
## Common Workflows
### Quality Assurance Workflow
Filter for completions with low evaluation scores on critical evaluation criteria. Add an automation to route them to a review queue for further human review. Set up a second automation to notify your team when the volume of low-scoring completions exceeds a threshold, indicating a potential system issue.
### Golden Dataset Building
Filter for human-reviewed completions that have high evaluation scores and represent successful interactions. Automatically add them to your golden dataset for regression testing. This ensures your test suite continuously grows with validated, real-world examples.
### Guardrail Monitoring
Filter for completions where guardrails were triggered (such as PII detection, toxicity filters, or policy violations). Add an automation to route these to a review queue so your team can analyze why guardrails fired. Set up notifications to alert your security or compliance team when guardrail triggers spike, indicating potential issues.
### Cost Optimization
Filter for high-cost agent traces that use excessive tokens. Run additional evaluations on these to assess whether the quality justifies the cost. Route borderline cases to a review queue so your team can determine if optimizations are needed.
***
# Observability
Source: https://docs.freeplay.ai/core-concepts/observability/observability-dashboard
Monitor, debug, and optimize your LLM systems in production with comprehensive observability tools.
## Overview
Freeplay's Observability dashboard is your window into understanding what's actually happening in your LLM systems. As the foundation of the **Monitor** phase in Freeplay's data flywheel, it captures production data including LLM completions, customer feedback, and system events and that feeds that information directly into your improvement cycle.
By surfacing patterns and issues through evaluations and metrics, Observability helps you identify what to experiment and test next, creating a continuous loop of monitoring, analysis, and improvement that makes your AI application better with every interaction.
## Key Capabilities
### Multi-Level Monitoring
Track your LLM interactions at multiple levels of granularity. Monitor individual completions to see full request/response details for each LLM call. Analyze traces to understand multi-step agent workflows and tool interactions. View sessions for aggregated user interactions and complete conversation flows. This hierarchical structure lets you zoom in and out as needed, from high-level patterns down to specific interactions. Learn more about the relationship between sessions, traces, and completions.
### Performance Metrics & Visualization
Track critical metrics through interactive charts and graphs that reveal cost trends across models and prompts, latency patterns that highlight performance bottlenecks, and usage volumes that show traffic patterns and peak periods. View evaluation results including model-graded evals, human labels, and custom metrics all in one place. Toggle between daily and weekly views to identify both immediate issues and long-term trends, with charts automatically updating as new data flows in for real-time visibility into system health.
### Advanced Search
Create powerful custom searches to focus on specific aspects of your system by filtering on prompt templates, model versions, environments, metadata fields, input content, or output patterns. Combine multiple filters for complex queries to isolate exactly the data you need. Save frequently-used searches for quick access and share filter configurations with team members to ensure everyone can monitor the same critical scenarios. This targeted approach helps you cut through noise to monitor what matters most for your application.
## Working with the Dashboard
### Navigating Your Data
The Observability dashboard combines visual analytics with detailed logs to give you complete visibility into your system. Graphs reveal performance trends and patterns over time, while the table view lets you drill into individual completions, traces, or sessions for debugging. Together, these views help you quickly identify issues at a high level, then dive deep into specific interactions to understand root causes.
### Creating Actionable Insights
From any view in Observability, you can flag interesting examples to add to review queues for team evaluation and labeling. Build test datasets directly from production data to validate improvements. Using evaluations, labels and information logging you can create a rich understanding of your system's behavior and be able to drill down to what matters.
### Team Collaboration
Freeplay's Observability dashboard enables collaboration by letting you share direct links to specific completions, traces, or sessions with team members. Create and share team-wide saved searches for common monitoring scenarios, ensuring everyone has access to the same views.
## Common Use Cases
### Production Monitoring
Set up saved searches to track error rates and failure patterns, monitor for performance degradation, and catch cost spikes or unusual usage patterns. Saved searches also help you quickly investigate customer-reported issues by filtering to specific time periods, users, or error types.
### Quality Assurance
Use Observability to regularly monitor your system's production outputs and keep an eye on the evaluation metrics over time. The platform can help identify cases that need improvement and track how prompt performance changes over time as you iterate on your system.
***
What's Next
* [Review Queues](/core-concepts/review-queues)
* [Datasets](/core-concepts/datasets/datasets)
# Sessions, Traces, and Completions
Source: https://docs.freeplay.ai/core-concepts/observability/sessions-traces-and-completions
Understand Freeplay's three-level hierarchy for organizing your AI application logs: sessions, traces, and completions.
When it comes to Observability in Freeplay, there are three related objects to understand. From low to high level abstractions, there are:
* **Completions:** Atomic LLM calls made up of a prompt and a response or output from a model.
* **Traces:** *Optionally* used to group related completions, e.g. when multiple completions are used to generate chat turn or single agent flow
* **Sessions:** The container for all completions and traces that make up a customer interaction. These can be 1:1 with completions for a simple feature that just uses one prompt, or they can be very large at times e.g. an entire conversation thread between a single user and a chatbot over multiple hours.
# Completions
The core, atomic unit for observability in Freeplay is a completion. A completion represents a single call to an LLM provider.
In Freeplay each completion is tied to a single prompt template version.
Most LLM applications involve a collection of completions to deliver a single outcome or user experience. This is where Sessions and Traces come into play: They provide a way to group completions in whatever way makes sense for your application logic.
# Traces
Traces are the next level in the hierarchy. They can be *optionally* used to structure and visualize the logical behavior of your application within a Session. We do not recommend using traces for simple applications, e.g. if your user experience relies on running a single prompt one time.
Traces are helpful when you need a more fine-grained way to group completions within a Session. A trace can contain one or more completions, and a session can contain one or more traces.
In some applications, multiple LLM calls (completions) are made to compute a single output for the user. Consider an example in which a system takes in a user question, generates a query to retrieve data (Completion A), and then passes that data into another LLM call to generate the answer (Completion B). Traces allow you to bundle together completions A and B to express the relationship between them.
For more detail on implementation of Traces see our [SDK Docs](/freeplay-sdk#traces) and a [Full Code Recipe](/developer-resources/recipes/record-traces).
While all completions must be tied to a session, traces are entirely optional. You do not need to use them if they don't make sense in the context of your application logic.
Note that there is a special case for logging traces as part of [multi-turn chatbots](/practical-guides/multi-turn-chat-support) that lets you pass an `input_question` and `output_answer` to render the start and end of a chat turn between a user and a bot. Traces *must* be used to take advantage of this feature.
# Sessions
Sessions are the highest level organizing principle in Freeplay for Observability. A session can contain one or more completions that are logically related to each other. We create a session every time you write logs to Freeplay, even when there's only one completion.
The simplest example of grouping completions into a session might be multiple user questions in the context of a single chatbot conversation. Each question results in a completion — a single call to an LLM provider to answer the question. Each of those completions can be logged together in a session to denote their relationship to each other (i.e all being part of the same conversation thread).
For more detail on recording completions, traces or sessions see our [SDK Docs](/freeplay-sdk#record-an-llm-interaction).
## Sessions View
The Session view in Freeplay helps teams quickly identify and diagnose issues in their LLM applications. As shown in the image above from a chatbot session, the interface provides a comprehensive overview of everything that happened during a session:
1. Left sidebar navigation displays the hierarchical structure of your session, showing multiple traces with color-coded indicators that reveal evaluation performance at a glance (green for passing, red for failing evaluations)
2. Top-level session metrics show aggregated information including total cost (\$0.002759), token usage (5K input, 838 output), and overall evaluation results
3. Individual trace details present evaluation breakdowns (e.g., "7 High, 1 Low"), associated notes, query topics, and response quality assessments for each interaction in the conversation
4. Full context display shows the complete input and output for each trace, making it easy to understand what the user asked and how the system responded
This structured view allows teams to prioritize where to focus their attention. In the example shown, you can quickly see that one trace has a low evaluation score, with issues flagged in "Response Quality" (marked as "Too Short"). This makes it easy to identify which specific interaction in a multi-turn conversation needs investigation — all without leaving the session view. This approach makes debugging and quality assurance faster and more efficient.
**API Reference**: Key endpoints for working with observability data:
* **Completions**: [Record Completion](/developer-resources/api-reference#completions), [Search Completions](/developer-resources/api-reference#search-api)
* **Traces**: [Record Trace](/developer-resources/api-reference#traces), [Search Traces](/developer-resources/api-reference#search-api)
* **Sessions**: [List Sessions](/developer-resources/api-reference#sessions), [Search Sessions](/developer-resources/api-reference#search-api)
### What's next
Now that you understand how to build Sessions and Traces with Completions, explore these resources:
* [Multi-Turn Chatbot Support](/practical-guides/multi-turn-chat-support) - Implement conversation tracking
* [Agents](/practical-guides/agents) - Structure complex agent workflows
* [Glossary](/resources/glossary) - Definitions of key Freeplay terms
# Filtering and Search
Source: https://docs.freeplay.ai/core-concepts/observability/ui-filters
Understand how Freeplay's filters work to build effective searches across your observability data.
Freeplay's filtering system lets you search across millions of records quickly. Understanding how different field types are indexed and searched helps you build more effective filters and avoid unexpected results.
## How text search works
When you filter on text fields like inputs, outputs, or metadata values, Freeplay uses **tokenized phrase matching**. This is different from simple substring search and has important implications for how you construct your filters.
### Tokenization
Text is broken into individual words called **tokens**. During this process:
* Punctuation and special characters are removed
* Text is normalized (case is ignored)
* Words become searchable units
For example, the email `user@freeplay.ai` is tokenized into:
```
["user", "freeplay", "ai"]
```
The `@` and `.` characters are stripped and treated as word boundaries.
### Phrase matching
When you search, the tokens in your query must appear **adjacent and in order** in the indexed data. This is why the `contains` filter works differently than you might expect.
The `contains` filter searches for **complete tokens in sequence**, not arbitrary substrings. Searching for `free` will not match `freeplay` because `free` is not a complete token in the indexed text.
## Field types and their behavior
Different fields in Freeplay use different search behaviors depending on their data type:
| Type | Examples | Search Behavior |
| ---------------------- | ------------------------------------------------------------ | ----------------------------------------------------- |
| **Text fields** | Inputs, outputs, evaluation notes | Tokenized phrase matching via `contains` |
| **Categorical fields** | Model, provider, environment, prompt template, review status | Exact match selection |
| **Numeric fields** | Cost, latency, token counts | Range queries (greater than, less than, between) |
| **Key-value fields** | Custom metadata, feedback, trace inputs/outputs | Key matched exactly; value uses tokenized text search |
### Categorical fields
Fields like **model**, **provider**, **environment**, and **prompt template** use exact matching. You select from a predefined list of values, and only records with that exact value are returned. These filters are straightforward and don't have tokenization considerations.
### Numeric fields
Fields like **cost**, **latency**, and **token counts** support range queries. You can filter for values greater than, less than, equal to, or between specific numbers.
### Key-value fields
For structured data like **custom metadata** or **feedback**, the filter has two parts:
1. **Key name**: Matched exactly (e.g., `customer_email`)
2. **Value**: Uses tokenized text search
This means if you have metadata like `{"customer_email": "morgan@freeplay.ai"}`, you can:
* Filter on the exact key `customer_email`
* Search the value using tokenized matching (same rules as text fields)
## Understanding the "contains" filter
The `contains` filter is the most common source of confusion. Here's what it actually means:
`contains` means "contains these complete tokens in this order" — not "contains this substring anywhere."
### What works vs. what doesn't
Given the indexed value `user@freeplay.ai` (tokenized to `["user", "freeplay", "ai"]`):
| Search Query | Result | Why |
| ------------------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| `user` | Matches | Complete token |
| `freeplay` | Matches | Complete token |
| `ai` | Matches | Complete token |
| `morgan freeplay` | Matches | Tokens in correct order |
| `freeplay.ai` | Matches | Tokenizes to `["freeplay", "ai"]` which matches |
| `freeplay ai` | Matches | Same as above |
| `@freeplay.ai` | Matches | `@` is stripped; tokenizes to `["freeplay", "ai"]` |
| `user@freeplay.ai` | Matches | Full value always matches |
| `use` | No match | Partial token |
| `free` | No match | Partial token |
| `@freeplay` | No match | Tokenizes to `["freeplay"]` but searches require the token sequence; the `@` makes it look for a token boundary that doesn't align |
| `ai freeplay` | No match | Tokens in wrong order |
## Common search scenarios
### Email addresses
Email addresses are tokenized at the `@` and `.` characters.
**Example**: `someone@freeplay.ai` becomes `["someone", "freeplay", "ai"]`
**Searches that work:**
* `someone` - complete token
* `freeplay` - complete token
* `freeplay.ai` or `freeplay ai` - token sequence
* `someone@freeplay.ai` - full value
* `@freeplay.ai` - tokenizes to match `["freeplay", "ai"]`
**Searches that don't work:**
* `@freeplay` - doesn't match because the search pattern doesn't align with token boundaries
* `play` - partial token, not indexed separately
* `one` - partial token
### Identifiers with special characters
Special characters act as token boundaries, which can cause **unexpected matches**.
**Example**: The value `user-123` is tokenized to `["user", "123"]`
This means all of the following will match the same records:
* `user-123`
* `user 123`
* `user_123`
* `user/123`
They all tokenize to the same sequence: `["user", "123"]`
If you need to distinguish between `user-123` and `user_123`, tokenized search won't help. Consider using a categorical field or adding a separate identifier field with exact matching.
### Custom metadata
When filtering on custom metadata:
1. Select the metadata key (exact match)
2. Enter the value to search (tokenized matching)
**Example**: If you have sessions with `{"user_type": "premium-enterprise"}`:
* Key: `user_type` (must match exactly)
* Value search for `premium` will match
* Value search for `enterprise` will match
* Value search for `prem` will not match (partial token)
### Evaluation results and notes
Evaluation result values and note content also use tokenized search:
* Filter by evaluation name (categorical/exact match)
* Search within results or notes (tokenized text search)
## Tips for effective filtering
\*\*Log Useful Info: \*\*When recording metadata, feedback or additional info to Freeplay, break it down into parts that your team may need to search by. This can help teams find useful information quickly!
**Use complete tokens**: When searching text fields, use full words rather than partial strings. If you're looking for emails from a domain, search for the domain name (`freeplay`) rather than a partial match (`@freeplay`).
**Combine filters to narrow results**: If tokenization gives you too many matches, add additional filters (time range, environment, model) to narrow down results.
# Advanced Templating (Mustache)
Source: https://docs.freeplay.ai/core-concepts/prompt-management/advanced-prompt-templating-using-mustache
Use Mustache syntax to add conditionals, loops, and dynamic content to your prompt templates.
## Overview
In addition to the basic prompt template features in Freeplay, we also support advanced prompt templating based on Mustache ([full docs here](https://mustache.github.io/mustache.5.html)). Mustache is a logic-less template syntax used to dynamically render new data within your prompts. It can be used for more advanced prompt templating use cases, like using conditional statements, if/else logic, and passing lists of variables into a prompt.
Some benefits include:
* **Dynamic Content Generation:** Create more flexible and context-aware prompts.
* **Simplified Template Management:** Less need for multiple similar templates, as one template can handle various data scenarios.
In practice, Mustache gives advanced controls to create more complex prompt templates, and is still user-friendly enough for non-developers to interpret and edit in Freeplay. However, just like the variable names described above, using advanced Mustache features depends on a strong contract with the code that will be used by your system to interact with the prompt template. It's important to align on expectations about what's in the code for anyone editing prompts in Freeplay.
Basic Mustache Syntax
[Mustache](https://mustache.github.io/mustache.5.html) templates are simple yet powerful. They work by expanding tags in a template using values provided in a hash or object. A Mustache template might look like this:
```
{{#repo}} // IF
{{name}}
{{/repo}}
{{^repo}} // ELSE
No repos :(
{{/repo}}
```
Given a hash with data for `repo` and `name` Mustache will render this template with appropriate substitutions.
**Supported Tag Types**
* **Variables:** Basic tags like \{\{name}} are replaced with the value associated with that key.
* **Dotted Names:** For nested objects, use dotted names like \{\{client.name}}.
* **Sections:** For conditional rendering, use sections like \{\{#logged\_in}}...\{\{/logged\_in}}. This will render the block only if logged\_in is true or non-empty.
* **Inverted Sections:** Use \{\{^logged\_in}}...\{\{/logged\_in}} for the opposite: rendering when logged\_in is false or empty.
Note that we do not follow the Mustache spec exactly. We don't currently support all tag types (see limitations below). Also, our Mustache implementation does /not/ escape special characters (quotes, braces, etc.). We think this is more intuitive, but is not technically aligned to the Mustache spec.
**Limitations**
We currently disallow Mustache Partials (syntax like `{{> prompt_rules}}`) that would recursively include other prompts or code.
***
What's Next
Dig deeper into Prompt Management and learn about Prompt Bundling, or jump into Evaluations.
* [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling)
* [Evaluations](/core-concepts/evaluations/evaluations)
# Environments
Source: https://docs.freeplay.ai/core-concepts/prompt-management/deployment-environments
Manage prompt deployment across development, staging, and production environments.
# Environments
Managing prompt deployment across different stages of your development lifecycle.
Environments in Freeplay allow you to control how and where prompt versions are deployed throughout your development workflow. By configuring environments, you can safely test changes, manage multiple deployment stages, and maintain different versions of prompts for various use cases.
## Common Use Cases
Many teams make use of prompt bundling in production. See our guide on [Prompt
Bundling](/core-concepts/prompt-management/prompt-bundling) for reference.
### Staged Deployment
The most common use of environments is to create one for each environment that teams use. For example if your team follows a typical testing, staging, production environment flow, you would create a tag for each.
This provides you the ability to rapidly test and iterate with the "test" environment tag and then once a new prompt version is ready, you can promote it to staging and finally production.
### Shadow Testing
Create a tag for your main prompt and a secondary tag to run new or experimental versions to compare performance without affecting user experience. Log interactions from both versions to evaluate and monitor improvements before full deployment.
### Rapid Iteration
Teams commonly create a "sandbox" or similar testing environment that allows team members to rapidly iterate on prompts without impacting other environments. Team members can use the "latest" tag for this or create their own to prevent overlapping with others.
### Feature Flags & A/B Testing
Use environments to manage feature rollouts or run A/B tests. Deploy different prompt variants to specific environments and route traffic accordingly based on your testing strategy.
## Setting Up Environments
Navigate to Settings > Environments to view and configure your environments.
Freeplay includes three default environments:
* latest - Automatically assigned to newly created prompt versions. Ideal for initial testing and development work.
* dev - Development environment for active iteration and experimentation.
* prod - Your stable, live environment serving end users.
You can customize these environments or create additional ones to match your team's workflow.
## Deploying Prompts to Environments
When you create or edit a prompt template and save a new version, Freeplay automatically assigns it to the *latest* environment. To deploy a prompt version to a different environment:
1. Open the prompt template
2. Select the version you want to deploy
3. Choose your target environment & select Deploy
The next time this prompt is fetched from Freeplay it will use the current deployed version! Only one version of a prompt can be deployed to each environment at a time.
## Using Environments in Your Code
We recommend setting your environment tags as environment variables so they do not change for production deployments, while for quick testing you can easily set in code:
```python python theme={null}
# Fetch the latest version
formatted_prompt = fpClient.prompts.get_formatted(
project_id=project_id,
template_name="your-template-name",
environment="latest",
variables=prompt_vars
)
# Fetch the production version
formatted_prompt = fpClient.prompts.get_formatted(
project_id=project_id,
template_name="your-template-name",
environment=os.getenv("FREEPLAY_ENVIRONMENT"),
variables=prompt_vars
)
```
```typescript typescript theme={null}
// Fetch the latest version
const formattedPrompt = await fpClient.prompts.getFormatted({
projectId: projectId,
templateName: "your-template-name",
environment: "latest",
variables: promptVars
});
// Fetch the production version
const formattedPrompt = await fpClient.prompts.getFormatted({
projectId: projectId,
templateName: "your-template-name",
environment: process.env.FREEPLAY_ENVIRONMENT,
variables: promptVars
});
```
***
[Prompt Bundling](/core-concepts/prompt-management/prompt-bundling)
[Observability](/core-concepts/observability/observability-dashboard)
# Prompt Management
Source: https://docs.freeplay.ai/core-concepts/prompt-management/managing-prompts
Version, iterate, and deploy prompt templates with full history tracking and environment-based deployment.
Freeplay simplifies the process of iterating and testing different versions of your prompts and provides a comprehensive prompt management system.
Here’s how it works, starting with the concept of a "prompt template" and then explaining options for managing prompts via the Freeplay UI or from your code.
# Prompt Templates
### What is a Prompt Template?
A prompt template in Freeplay defines the configuration for an LLM interaction: the messages, model settings, and optional tools or output schemas. Each template can have multiple versions as you iterate and refine your prompts.
### Components of a Prompt Template Version
The following elements make up the configuration of a prompt template version.
* [Content](#content) (e.g. messages)
* [Model config](#model-config)
* [Tools](#tools-optional) (optional)
* [Structured output schemas](#structured-outputs-optional) (optional)
When you make a change to any of these elements, that is considered a new prompt template version since any of them will affect the model's output.
Create new prompt versions frequently as you iterate. Think of versioning like committing code—each meaningful change gets its own version. See [Deployment Environments](/core-concepts/prompt-management/deployment-environments) for how to promote versions across environments.
#### Content
This is the actual text content of your prompt. The constant parts of your prompt will be written as normal text while the variable parts of your prompt will be denoted with variable place holders. When the prompt is invoked, the application specific content will be injected into the variables at runtime.
Take the following example:
In this example there are some input variables denoted in mustache syntax: `question`, `conversation_history`, and `supporting_information`. When invoked the prompt will become fully hydrated and what is actually sent to the LLM would look like this
Variables in Freeplay are defined via mustache syntax, you can find more information on advanced mustache usage for things like conditionals and structured inputs in [this guide](/core-concepts/prompt-management/advanced-prompt-templating-using-mustache).
#### Model Config
Model configuration includes model selection as well as associated parameters like temperature, max tokens, etc.
#### Tools (optional)
Tool schemas can also be managed in Freeplay as part of your prompt templates. Just like with content messages and model configuration the tool schema configured in Freeplay will be passed down in the SDK to be used in code. See more details on working with tools in Freeplay [here](/practical-guides/tools).
#### Structured Output Schemas (optional)
Structured outputs allow you to rely on the model outputting its results in a consistent format. Many models support a JSON output mode that can be defined and they will respect that format. Additionally, some model providers support the ability to define typed outputs that are more structured and formal than JSON outputs. Freeplay supports both. You can learn more [here](/core-concepts/prompt-management/structured-outputs/structured-outputs).
# Prompt Management
Freeplay offers a number of different features related to prompt management including
* Native versioning with full version history for transparent traceability
* An interactive prompt editor equipped with dozens of different models and integrated with your dataset
* A deployment mechanism tied into the Freeplay SDK
* A structured templating language for writing prompts and associated evaluations
Given that prompts are such a critical component of any LLM system, it’s important that prompts are represented in Freeplay so they can provide structure for observability, evaluation, dataset management, and experimentation. However we recognize that teams will have different preferences for where the ultimate source of truth is for prompts.
**Freeplay supports two modes of operating:**
* Freeplay as the source of truth for prompts, or
* Code as the source of truth for prompts
Here’s how prompt management works in each case.
### Freeplay as the source of truth for prompts
In this usage pattern new prompts are created within Freeplay and then passed down in code. Freeplay becomes the source of truth for the most up to date version of a given prompt. The flow follows the steps below.
If you don't already have a prompt template set up you can create one by going to Prompts → Create Prompt Template.
If you do already have a prompt template set up you can start making changes to your prompt template and you'll see that prompt turn to an unsaved draft.
As you make edits you can run your prompt against your dataset examples in real-time to understand the impacts of your changes.
Once you are happy with your new version you can hit Save and a new prompt template version is created.
Freeplay enables you to deploy different prompt version across your various environments helping facilitate the traditional environment promotion flow.
By default a new prompt template version will be tagged with *latest.* Once you're ready to promote that prompt template version to other environments you can add additional environments by hitting the "Deploy" button.
We always recommend running tests before deploying to upper environments, more on that [here](/core-concepts/test-runs/component-level-test-runs#testing-via-ui)!
Once a prompt version has been created it can be fetched via the SDK to be used in code.
There a number of different ways to fetch prompt templates (full details [here](/freeplay-sdk#prompts)) but the most common method is to retrieve a formatted prompt from a specific environment
```python python theme={null}
formatted_prompt = fpClient.prompts.get_formatted(
project_id=project_id,
template_name="rag-qa",
environment="dev",
variables={"keyA": "valueA"}
)
```
This will retrieve the version of the rag-qa prompt that is tagged with the dev environment. It will also inject the proper variable values to form a fully hydrated prompt.
This returns a formatted prompt object which is a helpful data object to be used when calling your LLM. You can key off of the object for the messages and model information and know it will all be formatted properly for your provider.
```python python theme={null}
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
```
By default prompt retrieval will fetch prompts from the Freeplay server each time. To remove that dependency you can use [prompt bundling](/core-concepts/prompt-management/prompt-bundling) which will instead copy the prompts to your local filesystem and retrieve them from there removing the network call altogether.
When [recording back to Freeplay](/freeplay-sdk#record-an-llm-interaction) you will pass through the prompt information to associated the specific prompt version with that observed completion.
```python python theme={null}
from freeplay import RecordPayload
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted_prompt.prompt_info,
start_time=start,
end_time=end,
usage=UsageTokens(
chat_response.usage.prompt_tokens,
chat_response.usage.completion_tokens
)
),
)
# record the LLM interaction
fp_client.recordings.create(payload)
```
By linking the prompt to the observed completion you will get a structured recording like this.
You can even open the completion back in the prompt editor to rerun it and make further tweaks to continue the iteration cycle.
### Code as the source of truth for prompts
In this usage pattern prompts live in your codebase and you push new versions to Freeplay programmatically. This gives you:
* **Code-first workflow**: Any changes happen solely in your code
* **Standard review process**: Use your existing code review workflow for prompt changes
* **Full flexibility**: Programmatically manage all prompt aspects
* **Automatic sync**: Push prompt updates to Freeplay as part of your CI/CD pipeline
Source code becomes the source of truth, but prompt versions are reflected in Freeplay for observability, experimentation in our playground, and more powerful evaluation targeting.
You can also create templates in the [Freeplay UI](/getting-started/start-in-ui#1-create-your-first-prompt-template).
Freeplay provides APIs to create and version prompt templates programmatically. For simplicity, we recommend using the [Create prompt template version by name](/api-reference/prompt-templates/create-prompt-template-version-by-name) endpoint.
Use the `create_template_if_not_exists=true` query parameter to create a template if it doesn't exist yet, or add a new version if it does. This enables an "upsert" workflow ideal for CI/CD pipelines.
```python Python theme={null}
import os
import requests
FREEPLAY_API_KEY = os.getenv("FREEPLAY_API_KEY")
project_id = os.getenv("FREEPLAY_PROJECT_ID")
base_api_url = "https://app.freeplay.ai/api/v2"
template_name = "my-assistant"
def sync_prompt_to_freeplay():
# Use create_template_if_not_exists to create or update
url = f"{base_api_url}/projects/{project_id}/prompt-templates/name/{template_name}/versions?create_template_if_not_exists=true"
# Load your prompt content (from file, config, etc.)
prompt_template = {
"template_messages": [
{
"role": "system",
"content": "You are a helpful assistant. The user's name is {{user_name}}."
},
{
"role": "user",
"content": "{{user_input}}"
}
],
"provider": "openai",
"model": "gpt-4o",
"llm_parameters": {
"temperature": 0.2,
"max_tokens": 1024
},
"version_name": "v1.2.0",
"version_description": "Production release with improved system prompt"
}
response = requests.post(
url,
headers={
"Authorization": f"Bearer {FREEPLAY_API_KEY}",
"Content-Type": "application/json"
},
json=prompt_template
)
response.raise_for_status()
result = response.json()
print(f"Created template {result['prompt_template_id']} version {result['prompt_template_version_id']}")
return result
# Run during CI/CD or deployment
if __name__ == "__main__":
sync_prompt_to_freeplay()
```
```typescript TypeScript theme={null}
const FREEPLAY_API_KEY = process.env.FREEPLAY_API_KEY;
const projectId = process.env.FREEPLAY_PROJECT_ID;
const baseApiUrl = "https://app.freeplay.ai/api/v2";
const templateName = "my-assistant";
async function syncPromptToFreeplay() {
// Use create_template_if_not_exists to create or update
const url = `${baseApiUrl}/projects/${projectId}/prompt-templates/name/${templateName}/versions?create_template_if_not_exists=true`;
// Load your prompt content (from file, config, etc.)
const promptTemplate = {
template_messages: [
{
role: "system",
content: "You are a helpful assistant. The user's name is {{user_name}}."
},
{
role: "user",
content: "{{user_input}}"
}
],
provider: "openai",
model: "gpt-4o",
llm_parameters: {
temperature: 0.2,
max_tokens: 1024
},
version_name: "v1.2.0",
version_description: "Production release with improved system prompt"
};
const response = await fetch(url, {
method: "POST",
headers: {
Authorization: `Bearer ${FREEPLAY_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify(promptTemplate)
});
if (!response.ok) {
throw new Error(`HTTP error! status: ${response.status}`);
}
const result = await response.json();
console.log(`Created template ${result.prompt_template_id} version ${result.prompt_template_version_id}`);
return result;
}
// Run during CI/CD or deployment
syncPromptToFreeplay();
```
```bash cURL theme={null}
curl -X POST "https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/name/my-assistant/versions?create_template_if_not_exists=true" \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"template_messages": [
{
"role": "system",
"content": "You are a helpful assistant. The user'\''s name is {{user_name}}."
},
{
"role": "user",
"content": "{{user_input}}"
}
],
"provider": "openai",
"model": "gpt-4o",
"llm_parameters": {
"temperature": 0.2,
"max_tokens": 1024
},
"version_name": "v1.2.0",
"version_description": "Production release with improved system prompt"
}'
```
The API returns the `prompt_template_id` and `prompt_template_version_id` which you'll use when recording completions.
**CI/CD Integration**: Most teams sync prompts during their build or deployment process. This ensures Freeplay always reflects your production prompt versions. See the [Create prompt template version by name](/api-reference/prompt-templates/create-prompt-template-version-by-name) endpoint for full documentation.
When recording completions, include the prompt template version ID to link observability data to specific prompt versions:
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
from freeplay.resources.prompts import PromptVersionInfo
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api"
)
# Your prompt (defined in code)
messages = [
{"role": "system", "content": "You are a helpful assistant. The user's name is Alice."},
{"role": "user", "content": "Tell me about the weather."}
]
# Make LLM call
response = openai_client.chat.completions.create(
model="gpt-4o",
messages=messages
)
# Record with prompt version info
all_messages = messages + [
{"role": "assistant", "content": response.choices[0].message.content}
]
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs={"user_name": "Alice", "user_input": "Tell me about the weather."},
call_info=CallInfo(provider="openai", model="gpt-4o"),
prompt_version_info=PromptVersionInfo(
prompt_template_version_id=template_version_id, # This is returned from the prompt create step above
environment="production"
)
)
)
```
```typescript TypeScript theme={null}
import Freeplay from "freeplay";
const fpClient = new Freeplay({
freeplayApiKey: process.env.FREEPLAY_API_KEY,
baseUrl: "https://app.freeplay.ai/api"
});
// Your prompt (defined in code)
const messages = [
{ role: "system", content: "You are a helpful assistant. The user's name is Alice." },
{ role: "user", content: "Tell me about the weather." }
];
// Make LLM call
const response = await openaiClient.chat.completions.create({
model: "gpt-4o",
messages
});
// Record with prompt version info
const allMessages = [...messages, response.choices[0].message];
await fpClient.recordings.create({
projectId,
allMessages,
inputs: { userName: "Alice", userInput: "Tell me about the weather." },
callInfo: { provider: "openai", model: "gpt-4o" },
promptVersionInfo: {
promptTemplateVersionId: templateVersionId,
environment: "production"
}
});
```
By linking prompt template versions to related completions, you get structured logging, can target prompts with evaluations, track each prompt version's usage and quality metrics, and run copies of the prompt directly from the Freeplay UI.
Even when treating code as the source of truth, you can use Freeplay's prompt editor to experiment with changes, then sync those changes back to code. Use the [Get all prompt templates by environment](/api-reference/prompt-templates/get-all-prompt-templates-by-environment) endpoint to export prompts.
***
## Next steps
See complete integration patterns for fetching and using prompts
Learn more about evaluation types and alignment
Execute batch tests to compare prompt versions and measure quality
# Prompt Bundling
Source: https://docs.freeplay.ai/core-concepts/prompt-management/prompt-bundling
Bundle prompt templates into your deployment artifacts for increased resilience and release control.
Prompt Bundling enables you to fetch prompt templates once from the Freeplay server during your build process and then store them directly in the filesystem of your deployed artifact. You can then configure the Freeplay client to call prompt templates and model configuration directly from your code, rather than the Freeplay server.
## Advantages of Prompt Bundling
* Release Management Controls
* This approach to deployment provides full control over what's running in production, since all prompt & model iterations managed in Freeplay get treated like any other part of your code.
* Facilitates compliance obligations for peer review, use of established version control systems like GitHub, etc. A detailed guide on using prompt bundling with GitHub Actions is [here](/security-compliance/production-prompt-bundling-compliance-guard-rails).
* Increased Production Resilience
* By decoupling the fetching of prompt templates from Freeplay, you ensure your application can keep serving requests if Freeplay is unavailable.
* Unintended prompt changes cannot hurt production because prompts are only refreshed during your application build process.
* Reduced Latency
* By reading the prompt template from your application filesystem instead of fetching it from Freeplay on each request, you can slightly reduce latency overhead.
## Disadvantages of Prompt Bundling
* Increased Cycle Time from Experimentation to Production
* Requiring a build process to get new prompts into production, increasing cycle time between prompt updates and the deployment of those updates into production.
Many Freeplay customers use a hybrid approach: server-side prompt management in lower environments (dev/staging) for rapid iteration, and prompt bundling in production for release control. See [Multi-Environment Configuration](#multi-environment-configuration) below.
## Implementation
Downloading of prompts is handled by the Freeplay Python SDK
### Install the Python SDK
Python will be used to download the prompts for all SDKs
```bash theme={null}
pip install freeplay
```
### Download the Prompt Templates
```bash theme={null}
Usage: freeplay download [OPTIONS]
Options:
--project-id TEXT The Freeplay project ID. [required]
--environment TEXT The environment from which the prompts will be pulled.
[required]
--output-dir TEXT The directory where the prompts will be saved.
[required]
--help Show this message and exit.
```
```bash theme={null}
# Set necessary environment variables
export FREEPLAY_API_KEY=$FREEPLAY_API_KEY
export FREEPLAY_SUBDOMAIN=$FREEPLAY_SUBDOMAIN
export FREEPLAY_PROJECT_ID=$FREEPLAY_PROJECT_ID
# Run the command with your project and environment settings, specifying where you want the prompts to be placed.
freeplay download --project-id= --output-dir=my_freeplay_prompts --environment=prod
```
### Download all Prompt Templates
To download all prompt templates in your account, you can use the `download-all` option for the API. This will download all prompts and store them by project id. For private projects, prompts will only be downloaded if the user has access to that project.
```bash theme={null}
# Set necessary environment variables
export FREEPLAY_API_KEY=$FREEPLAY_API_KEY
export FREEPLAY_SUBDOMAIN=$FREEPLAY_SUBDOMAIN
# Run the command with your project and environment settings, specifying where you want the prompts to be placed.
freeplay download-all --output-dir=my_freeplay_prompts --environment=prod
```
### Point your Freeplay Client at your Prompt Directory
```python python theme={null}
from freeplay import Freeplay
from freeplay.resources.prompts import FilesystemTemplateResolver
from pathlib import Path
# create a freeplay client link to local filesystem
fpClientLocal = Freeplay(
freeplay_api_key=freeplay_key,
api_base=freeplay_api_base,
template_resolver=FilesystemTemplateResolver(Path(freeplay_template_path))
)
```
```typescript typescript theme={null}
import Freeplay, { FilesystemTemplateResolver } from "freeplay";
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: process.env["FREEPLAY_URL"],
templateResolver: new FilesystemTemplateResolver(templatePath),
});
```
```java java theme={null}
// create the client
Path templateDir = Paths.get(outputDirectory);
Freeplay localClient = new Freeplay(Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(freeplaySubdomain)
.templateResolver(new FilesystemTemplateResolver(templateDir))
.providerConfigs(new ProviderConfigs(new AnthropicProviderConfig(anthropicApiKey)))
);
```
## Multi-Environment Configuration
A common pattern is to use different Freeplay client configurations for different environments. This approach gives you the best of both worlds:
* **Dev/Staging**: Fetch prompts from the Freeplay server for rapid iteration and experimentation
* **Production**: Use prompt bundling for resilience and release control
```python python theme={null}
from freeplay import Freeplay
from freeplay.resources.prompts import FilesystemTemplateResolver
from pathlib import Path
import os
def create_freeplay_client():
"""Create a Freeplay client configured for the current environment."""
base_config = {
"freeplay_api_key": os.environ["FREEPLAY_API_KEY"],
"api_base": os.environ["FREEPLAY_URL"],
}
if os.environ.get("ENVIRONMENT") == "production":
# Production: Use bundled prompts from filesystem
return Freeplay(
**base_config,
template_resolver=FilesystemTemplateResolver(
Path(os.environ["FREEPLAY_TEMPLATE_PATH"])
)
)
else:
# Dev/Staging: Fetch prompts from Freeplay server
return Freeplay(**base_config)
fp_client = create_freeplay_client()
```
```typescript typescript theme={null}
import Freeplay, { FilesystemTemplateResolver } from "freeplay";
function createFreeplayClient(): Freeplay {
const baseConfig = {
freeplayApiKey: process.env.FREEPLAY_API_KEY!,
baseUrl: process.env.FREEPLAY_URL!,
};
if (process.env.ENVIRONMENT === "production") {
// Production: Use bundled prompts from filesystem
return new Freeplay({
...baseConfig,
templateResolver: new FilesystemTemplateResolver(
process.env.FREEPLAY_TEMPLATE_PATH!
),
});
} else {
// Dev/Staging: Fetch prompts from Freeplay server
return new Freeplay(baseConfig);
}
}
const fpClient = createFreeplayClient();
```
```java java theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.Config;
import ai.freeplay.client.thin.resources.prompts.FilesystemTemplateResolver;
import java.nio.file.Paths;
public class FreeplayClientFactory {
public static Freeplay create() {
Config config = new Config()
.freeplayAPIKey(System.getenv("FREEPLAY_API_KEY"))
.customerDomain(System.getenv("FREEPLAY_SUBDOMAIN"));
if ("production".equals(System.getenv("ENVIRONMENT"))) {
// Production: Use bundled prompts from filesystem
config.templateResolver(new FilesystemTemplateResolver(
Paths.get(System.getenv("FREEPLAY_TEMPLATE_PATH"))
));
}
// Dev/Staging: Uses default server-side resolution
return new Freeplay(config);
}
}
```
This pattern allows your team to iterate quickly on prompts in lower environments while maintaining strict control over what's deployed to production.
***
## What's Next
Learn how to use Mustache syntax for advanced prompt templating or move onto the next section on Evaluations.
* [Advanced Prompt Templating Using Mustache](/core-concepts/prompt-management/advanced-prompt-templating-using-mustache)
* [Evaluations](/core-concepts/evaluations/evaluations)
# JSON Mode
Source: https://docs.freeplay.ai/core-concepts/prompt-management/structured-outputs/json-mode
Get valid JSON output from LLMs without enforcing a strict schema.
## When to use JSON mode
JSON mode is ideal for:
* **Flexible data structures**: When the exact output structure may vary based on input
* **Rapid prototyping**: Quick experimentation without defining formal schemas
* **Simple JSON needs**: Cases where valid JSON is sufficient without strict validation
**When to use structured outputs instead:**
* Production applications requiring guaranteed schema compliance
* Data that feeds directly into typed systems or databases
* Cases where validation and retry logic would be complex
See the [Structured Outputs documentation](/core-concepts/prompt-management/structured-outputs/structured-outputs-openai) for schema-enforced JSON output.
## Enabling JSON mode in Freeplay
### In the prompt editor
1. Open your prompt template in the Freeplay editor
2. Navigate to the output configuration section
3. Select "Enable JSON output"
4. Save your prompt template
### In your prompts
When using JSON mode, always include explicit instructions in your prompt to output JSON:
```
You are a helpful assistant that outputs JSON.
Extract the key information from the following text and return it as JSON.
Include fields for name, date, and main topics discussed.
Your response should look like:
{
"name": ,
"date": ,
"topics":
}
```
**Important**: The API will return an error if the string "JSON" doesn't appear somewhere in your messages when JSON mode is enabled.
## Best practices
### Always instruct the model
Include clear instructions in your system message or prompt:
```
Return your response as valid JSON with the following structure:
{
"summary": "brief summary here",
"key_points": ["point 1", "point 2"],
"sentiment": "positive/negative/neutral"
}
```
### Validate the output
JSON mode guarantees valid JSON syntax, but not a specific structure. Always validate the output:
```python python theme={null}
import json
try:
data = json.loads(response.choices[0].message.content)
# Validate expected fields
if "summary" not in data:
# Handle missing fields
pass
except json.JSONDecodeError: # Handle parse errors (rare with JSON mode)
pass
```
### Provide examples
For consistent output structures, include examples in your prompt:
```
Example output format:
{
"title": "Meeting Notes",
"date": "2024-10-15",
"attendees": ["Alice", "Bob"],
"action_items": [
{"task": "Review document", "owner": "Alice", "due": "2024-10-20"}
]
}
```
### Handle edge cases
JSON mode can fail in certain edge cases:
* **Token limit reached**: Output may be incomplete JSON
* **Model refuses**: Safety refusals may not be valid JSON
Always check `finish_reason`:
```python python theme={null}
if response.choices[0].finish_reason != "stop":
# Handle incomplete response
print(f"Response incomplete: {response.choices[0].finish_reason}")
```
## Recording to Freeplay
Record your JSON mode completions to Freeplay like any other completion:
**Python:**
```python python theme={null}
from freeplay import RecordPayload, CallInfo
# Record the completion
session = fp_client.sessions.create()
payload = RecordPayload(
project_id=project_id,
all_messages=[
\*formatted_prompt.llm_prompt,
{
"role": response.choices[0].message.role,
"content": response.choices[0].message.content,
}
],
inputs={"text": user_input},
session_info=session.session_info,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
)
fp_client.recordings.create(payload)
```
## JSON mode vs structured outputs
| Feature | JSON Mode | Structured Outputs |
| -------------------------- | -------------------------- | -------------------------- |
| Valid JSON guaranteed | ✓ | ✓ |
| Schema compliance | ✗ | ✓ |
| Schema definition required | ✗ | ✓ |
| Flexibility | High | Lower (strict schema) |
| Model support | Broad | OpenAI gpt-4o+ only |
| Validation needed | ✓ (manual) | ✗ (automatic) |
| Best for | Flexible JSON, prototyping | Production, strict schemas |
**Recommendation**: Use structured outputs for production applications requiring reliable, schema-compliant data. Use JSON mode for prototyping or when you need flexible JSON without strict validation.
## Troubleshooting
**Issue**: Model not outputting JSON
* **Solution**: Ensure "JSON" appears in your prompt instructions. Add explicit JSON format instructions.
**Issue**: Incomplete JSON output
* **Solution**: Check `finish_reason`. If it's `length`, increase `max_tokens`. If it's `content_filter`, adjust your input.
**Issue**: Inconsistent output structure
* **Solution**: Provide clear examples in your prompt. Consider using structured outputs if you need guaranteed schema compliance.
***
[Structured Outputs (OpenAI)](/core-concepts/prompt-management/structured-outputs/structured-outputs-openai)
[Prompt Bundling](/core-concepts/prompt-management/prompt-bundling)
# Structured Outputs
Source: https://docs.freeplay.ai/core-concepts/prompt-management/structured-outputs/structured-outputs
Constrain LLM responses to specific formats for reliable parsing and type-safe integrations.
# Overview
The use of structured outputs help you control how your LLMs response will be formatted. You can configure the output in Freeplay and test it in the prompt playground.
## Why use Structured output?
Structured outputs provide consistency and safety to your AI outputs. This allows you to easily parse the output information in a consistent and expected manner. This can help eliminate parsing, reducing errors, add type safety and make your integrations reliable and safe for use.
## Supported Types
Freeplay supports two ways to get structured output from your LLMs:
### OpenAI Structured Outputs
**Guarantees schema compliance** - The LLM will always match your exact JSON schema.
**Best for:**
* Production applications requiring reliable, consistent data structures
* Data feeding into typed systems or databases
* Extracting structured information from unstructured data
**Supported models:** OpenAI gpt-4o and later
[**Learn more about Structured Outputs →**](/core-concepts/prompt-management/structured-outputs/structured-outputs-openai)
***
### JSON Mode
**Ensures valid JSON** - The LLM will output valid JSON without enforcing a specific structure.
**Best for:**
* Rapid prototyping and experimentation
* When you need JSON but don't require strict validation
* Working with models that don't support structured outputs
When using JSON mode, you must always instruct the model to produce JSON via some message in the conversation, for example via your system message. If you don't include an explicit instruction to generate JSON, the model may generate an unending stream of whitespace and the request may run continually until it reaches the token limit. To help ensure you don't forget, the API will throw an error if the string "JSON" does not appear somewhere in the context.
JSON mode will not guarantee the output matches any specific schema, only that it is valid and parses without errors. You should use Structured Outputs to ensure it matches your schema, or if that is not possible, you should use a validation library and potentially retries to ensure that the output matches your desired schema.
[**Learn more about JSON Mode →**](/core-concepts/prompt-management/structured-outputs/json-mode)
***
## Comparison
| Feature | Structured Outputs | JSON Mode |
| ------------------------------ | ------------------------------------ | ---------------------------- |
| **Valid JSON output** | ✓ Always | ✓ Always |
| **Schema compliance** | ✓ Guaranteed | ✗ Not guaranteed |
| **Schema definition required** | ✓ Yes (ex: Pydantic/Zod) | ✗ No |
| **Flexibility** | Lower (strict schema) | High |
| **Validation needed** | ✗ Automatic | ✓ Manual |
| **Model support** | OpenAI gpt-4o+ | Most OpenAI models |
| **Supported providers** | OpenAI, Azure OpenAI | OpenAI, Azure OpenAI |
| **Best for** | Production apps | Prototyping, flexible output |
| **Error handling** | Refusals programmatically detectable | Manual validation required |
## Quick decision guide
**Use Structured Outputs if:**
* You're building a production application
* You need guaranteed field presence and types
* Data feeds directly into databases or typed systems
* You want to eliminate validation logic and retries
**Use JSON Mode if:**
* You're prototyping or experimenting
* You need flexibility over strict compliance
* You want JSON output without defining formal schemas
## How Freeplay helps
Regardless of which approach you choose, Freeplay supports you by allowing you to quickly test and iterate against your output requirements in the prompt playground, render the structured outputs for easy review and analysis and Freeplay also allows you to record your structure defined in code and save it to your template for use.
***
What's Next
* [Structured Outputs (OpenAI)](/core-concepts/prompt-management/structured-outputs/structured-outputs-openai)
* [JSON Mode](/core-concepts/prompt-management/structured-outputs/json-mode)
# Structured Outputs (OpenAI)
Source: https://docs.freeplay.ai/core-concepts/prompt-management/structured-outputs/structured-outputs-openai
Use OpenAI's strict JSON schema mode to guarantee responses match your expected data structure.
## Key points
* **Define schemas in code**: Use a json schema such as Pydantic (Python) or Zod (Node.js) to define your schemas as the source of truth
* **Pass to Freeplay on record**: Include `output_schema` (in JSON format) in your `RecordPayload` when recording completions
* **Freeplay interprets automatically**: Freeplay extracts and stores your schema for use in testing and observability
* **OpenAI format**: OpenAI requires `response_format` with type `json_schema`, setting `strict: true`, and providing the schema
* **Keep code as source of truth**: While you can view and edit schemas in the Freeplay UI, we recommend maintaining schemas in your codebase for consistency
## How does Freeplay help with structured outputs?
Freeplay supports the complete lifecycle of working with structured outputs - by interpreting schemas passed via code and saving them with prompt templates to allowing you to view and manage output schemas alongside your prompt templates.
The SDK supports recording structured outputs and their associated schemas which are interpreted by Freeplay and can be associated with your prompt template. With the Freeplay web app, you can define output schemas alongside your prompt templates.
## Managing your output schema with Freeplay
To get started from your code, simply have your predefined schemas, see for example using Pydantic and Zod to define a very specific output:
```python python theme={null}
from pydantic import BaseModel
from typing import List
class ResumeWorkExperience(BaseModel):
company: str
position: str
start_date: str
end_date: str
description: str
class ResumeData(BaseModel):
full_name: str
email: str
phone: str
skills: List[str]
work_experience: List[ResumeWorkExperience]
```
```typescript typescript theme={null}
import { z } from "zod";
const ResumeWorkExperienceSchema = z.object({
company: z.string(),
position: z.string(),
startDate: z.string(),
endDate: z.string(),
description: z.string(),
});
const ResumeDataSchema = z.object({
fullName: z.string(),
email: z.string(),
phone: z.string(),
skills: z.array(z.string()),
workExperience: z.array(ResumeWorkExperienceSchema),
});
```
### Pass schemas to Freeplay when recording
When you record an LLM interaction to Freeplay, include the `output_schema` parameter. Freeplay will automatically interpret and store your schema, making it available in the UI for testing and iteration.
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
# After making your LLM call with structured outputs...
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session.session_info,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
output_schema=ResumeData.model_json_schema() # Pass your Pydantic schema
)
fp_client.recordings.create(payload)
```
```typescript typescript theme={null}
import { z } from "zod";
// After making your LLM call with structured outputs...
await fpClient.recordings.create({
projectId,
allMessages: messages,
sessionInfo: getSessionInfo(session),
inputs: inputVariables,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: callInfo,
outputSchema: z.toJSONSchema(ResumeDataSchema) // Pass your Zod schema
});
```
That's it! Freeplay now has your structured output schema and you can leverage it throughout the platform.
### Open a completion in the prompt editor
Once you have recorded a completion with your schema, you can have it automatically added to your Freeplay prompt template by opening the example in the prompt editor. This is when Freeplay will automatically interpret the structure and associate it with your prompt.
Now save this version of your prompt template and you are all done! You can now access the model parameter `formatted_prompt.formatted_output_schema` in your code to access the schema.
### Benefits of this approach
By defining schemas in code and passing them to Freeplay, you get:
* **Source of truth in code**: Your application code remains the authoritative definition of your schemas, ensuring consistency between development and production
* **Type safety**: Leverage your language's type system and validation libraries
* **Easy testing**: Run tests in Freeplay UI that respect your structured output format
* **Quick iteration**: Test prompt variations in Freeplay while maintaining schema compliance
* **Observability**: View structured outputs in the Freeplay dashboard with full schema context
## Testing with structured outputs
Once your schema is recorded in Freeplay, you can run tests from the UI that will respect your structured output format. This makes it easy to:
* Validate that prompt changes maintain schema compliance
* Test against datasets to ensure consistent output structure
* Debug schema validation issues before they reach production
* Compare different prompt versions while enforcing the same output structure
When you run a test from the Freeplay UI on a prompt with an associated structured output schema, the test will automatically use that schema to validate responses.
## Viewing schemas in the Freeplay UI
You can view and inspect structured output schemas in the Freeplay UI, though we recommend keeping your code as the source of truth for schema definitions.
When viewing a completion or prompt template in Freeplay, you'll see the associated output schema displayed. This is helpful for:
* Understanding what structure was used for a particular completion
* Reviewing schema definitions when debugging issues
* Sharing schema information with team members
While it's possible to define or modify schemas directly in the Freeplay UI, we recommend using this primarily for exploration and testing. For production use, maintain your schemas in code to ensure consistency and version control.
## Complete code examples
See a complete end to end example [here](/developer-resources/recipes/structured-outputs).
## Important notes
* **OpenAI only**: At this time, structured outputs are only supported with OpenAI models. Support for additional providers is coming soon.
* **Schema normalization**: Freeplay automatically adds `"additionalProperties": false` to all object types in your schema. This ensures the LLM doesn't add unexpected fields to your structured output.
* **Validation**: When you add an output schema to a prompt template, Freeplay validates that your selected model supports structured outputs. You'll receive an error if you try to use structured outputs with an unsupported model.
* **Strict mode**: OpenAI's structured outputs require `strict: true` in the `json_schema` configuration. This enforces exact schema compliance.
* **Nested schemas**: You can use nested objects and arrays in your schemas for complex structured outputs.
* **Optional vs required fields**: Use the `required` array to specify which fields must be present. Fields not in the `required` array are optional.
* **Programmatic schemas**: You can define schemas either in your prompt templates or programmatically using JSON schemas such as Pydantic (Python) or Zod (Node.js). Both approaches work with Freeplay's recording and observability.
***
[Structured Outputs](/core-concepts/prompt-management/structured-outputs/structured-outputs)
[JSON Mode](/core-concepts/prompt-management/structured-outputs/json-mode)
# Review Queues
Source: https://docs.freeplay.ai/core-concepts/review-queues
Streamline human review of your AI outputs with organized queues, team assignments, and actionable insights.
## Introduction
Streamline human review of your LLM systems with Freeplay's Review Queues. Organize work assignments, generate actionable insights, and transform observations into meaningful improvements.
Review Queues help you establish systematic, organized human review processes for your LLM systems in Freeplay. Easily select completions for your team to evaluate, assign specific work to team members, and generate comprehensive insight reports to share findings across your organization. Review Queues make it simple to maintain consistent review practices and transform observations into actionable improvements.
## Creating Review Queues
From the Reviews tab, you can create a new Review Queue with customized instructions for your team and an optional due date to help manage project timelines effectively.
## Building Effective Review Queues
Review Queues serve multiple purposes: addressing issues flagged by evaluations, investigating team-identified concerns, or conducting systematic quality assessments. By creating thoughtfully curated review queues, you gain the ability to dive deep into your data, understand underlying patterns, identify areas for improvement, and drive meaningful product enhancements.
Teams that actively use Review Queues develop a much deeper understanding of their LLM applications, enabling them to implement targeted changes that deliver real improvements to their products.
### Adding Data to Review Queues
You can add Traces or Completions to a Review Queue for comprehensive analysis. Choose from two convenient methods:
Bulk AddIndividual Add
Perfect for adding multiple items at once:
1. **Search** across specific criteria such as metadata, evaluation scores, or input fields
2. **Select** multiple Completions from the filtered results
3. **Add** them all to your Review Queue in one action
### Managing Review Workflows
Once you've created your Review Queue, you can assign specific Completions to different team members based on their expertise and availability. This ensures the right people are reviewing the most relevant content.
Domain experts and analysts can efficiently work through their assigned Completions, providing detailed feedback and analysis. The interface is designed to streamline the review process while capturing comprehensive insights.
## Generating Actionable Insights
Every Review Queue automatically generates a comprehensive insights report, enabling analysts to:
* **Share findings** broadly across the organization
* **Highlight key takeaways** for stakeholders
* **Identify actionable improvements** for the LLM system
* **Track progress** over time
These insights reports transform individual review observations into strategic recommendations that drive meaningful system improvements.
***
[Auto-categorization](/core-concepts/evaluations/auto-categorization)
[Datasets](/core-concepts/datasets/datasets)
# Component Level Test Runs
Source: https://docs.freeplay.ai/core-concepts/test-runs/component-level-test-runs
Test individual prompts and models in isolation for rapid iteration.
# Introduction
Component test runs focus on testing individual prompts and models in isolation. This targeted approach enables rapid iteration during development and makes testing accessible to all team members, regardless of technical expertise.
## Why Component Testing Matters
During prompt development, you need fast feedback on changes without the overhead of running your entire system. Component testing provides this rapid iteration cycle, allowing you to test dozens of variations in minutes rather than hours. It also democratizes testing—product managers, domain experts, and other stakeholders can validate outputs through the UI without writing code.
Component tests excel at isolating variables. When you change a prompt's instructions or switch models, you can see exactly how that specific change affects performance without other system components adding noise to your results.
## Testing via UI
The Freeplay UI provides no-code testing that's perfect for rapid prompt iteration. Navigate to your prompt template, select the version you want to test, and click the Test button. You'll configure your test by selecting a dataset, choosing model parameters, and naming your test run.
Once running, Freeplay processes each test case automatically—fetching inputs from your dataset, formatting the prompt with variables, sending requests to the model, and recording outputs with evaluations. Results appear in real-time as the test progresses.
The results page provides an overview of how the components behave. This view shows the scores of the different versions being tested. We highlight the better performers to help track which result has done better overall:
The results view provides row-level details, allowing you to dive in and see the granular changes for your component. In this view you can also mark your preference and see which component change has performed best:
Green highlights indicate improvements while gray shows regressions. The visual comparison makes it easy to spot patterns and decide which version performs best. You can add additional comparisons to test against other versions or previous test runs, building a comprehensive view of how your prompts evolve.
## Running Component Tests via SDK
While the UI excels at accessibility, the SDK provides programmatic control for automation and integration:
```python python theme={null}
from freeplay import Freeplay, RecordPayload
from openai import OpenAI
# Create test run
test_run = fp_client.test_runs.create(
project_id=project_id,
testlist="Golden Set",
name="Customer Support Prompt v2.3"
)
# Get prompt template
template_prompt = fp_client.prompts.get(
project_id=project_id,
template_name="customer-support",
environment="latest"
)
# Test each case
for test_case in test_run.test_cases:
formatted_prompt = template_prompt.bind(test_case.variables).format()
# Make LLM call
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
# Record results
fp_client.recordings.create(
RecordPayload(
all_messages=all_messages,
inputs=test_case.variables,
test_run_info=test_run.get_test_run_info(test_case.id),
# ... additional parameters
)
)
```
The SDK approach enables testing multiple models or parameter variations programmatically, integrating with CI/CD pipelines, and building custom testing workflows.
## Analyzing Results
Component tests provide detailed insights into prompt performance. Evaluation scores show how well your prompts meet quality criteria like correctness, safety, and formatting. Cost and latency metrics help you balance quality with performance requirements.
The row-level view lets you examine specific examples to understand failure modes. Click any test case to see the full input, output, and evaluation explanations. This granular inspection helps you identify patterns—perhaps your prompt struggles with certain input types or consistently fails specific evaluations.
## Best Practices
Start your testing with small datasets of 10-20 carefully chosen examples. This allows quick iteration while still catching major issues. Once you've refined your prompt, expand to larger datasets for comprehensive validation.
Version your prompts thoughtfully. Create new versions for significant changes, add clear descriptions explaining what changed, and maintain a changelog of iterations. This discipline helps you track what works and roll back if needed.
Use consistent datasets when comparing versions to ensure fair comparisons. A "Golden Set" of ideal examples serves as your north star, while edge case datasets help ensure robustness.
Establish clear quality thresholds before testing. Decide what evaluation scores indicate success, which metrics matter most for your use case, and when a prompt is ready for production. These standards guide your iteration and prevent endless tweaking.
## Common Testing Patterns
**A/B Testing Instructions**: Test variations of your prompt instructions to find the most effective phrasing. Keep everything else constant while changing specific instructions to isolate their impact.
**Model Comparison**: Run the same prompt across different models to find the best balance of quality, cost, and speed. This helps you choose between GPT-4's power and Claude's efficiency for your specific use case.
**Parameter Optimization**: Test different temperature settings, token limits, and other parameters to fine-tune behavior. Lower temperatures provide consistency while higher values enable creativity.
**Format Validation**: Ensure your prompts consistently produce the expected output format, whether that's JSON, markdown, or structured text. Add format checking evaluations to catch deviations early.
## Integration with Development Workflow
Component testing fits naturally into your prompt development cycle. Begin by testing your current production version to establish a baseline. Identify improvement opportunities from failed test cases or user feedback. Modify your prompt based on these insights, run component tests to validate changes, and compare results against your baseline.
Once tests pass your quality thresholds, you can deploy with confidence. Continue monitoring production performance and add new test cases based on real-world edge cases you discover.
***
[End-to-End Test Runs](/core-concepts/test-runs/end-to-end-test-runs)
[SDK Reference](/freeplay-sdk)
# End-to-End Test Runs
Source: https://docs.freeplay.ai/core-concepts/test-runs/end-to-end-test-runs
Test your complete AI system including agents, RAG pipelines, and multi-step workflows.
# Introduction
End-to-end test runs validate your entire AI system by passing test cases through your complete pipeline. This comprehensive approach ensures that changes to any component don't cause unexpected regressions elsewhere in your system.
## Why End-to-End Testing Matters
Modern AI applications consist of multiple interacting components—LLM calls in sequence, tool usage, retrieval systems, and agent orchestration. Testing individual pieces in isolation isn't enough. You need to understand how changes ripple through your entire system to catch issues before they reach users.
End-to-end tests provide realistic performance assessment by testing your system exactly as users experience it. They capture complex workflows including multi-step processes, tool usage, and agent decision-making while tracking both final outputs and intermediate steps.
## Implementation
End-to-end tests execute through the SDK, giving you complete control over your system's execution. Here's how to test a support agent system that uses multiple sub-agents and tools. This example is using Freeplay's Support Agent that helps us take in customer requests and make sure we are tracking them well. It is made up of several components including `FreeplaySupportAgent`, `DocsAgent` and a `LinearAgent`. Each of these agents handles different tasks and follow the common router prompt format for testing. We are using an [Agent (trace dataset)](/practical-guides/agents#datasets) in Freeplay to test the end to end behavior.
### Step 1: Set up
```python python theme={null}
import os
import time
from typing import Optional
from tqdm import tqdm
from openai import OpenAI
from freeplay import (
Freeplay,
RecordPayload,
SessionInfo,
TraceInfo,
TestRunInfo,
CallInfo,
)
from dotenv import load_dotenv
load_dotenv(override=True)
# Optional SDK helpers (present in recent Freeplay SDKs)
try:
from freeplay import UsageTokens # type: ignore
except Exception:
pass
UsageTokens = None
# TODO: Update these to your environment variables
FREEPLAY_API_KEY = os.environ.get("FREEPLAY_API_KEY") or ""
OPENAI_API_KEY = os.environ.get("OPENAI_API_KEY") or ""
if not FREEPLAY_API_KEY:
raise RuntimeError("FREEPLAY_API_KEY is not set.")
if not OPENAI_API_KEY:
raise RuntimeError("OPENAI_API_KEY is not set.")
# TODO: Update these to your configuration
PROJECT_ID = "" # TODO: Update this to your project ID
TRACE_DATASET_NAME = "" # TODO: Update this to your dataset name that targets an agent
TEST_RUN_NAME = "" # TODO: Update this to your test name
TEMPLATE_NAME = "" # TODO: Update this to your prompt name
TEMPLATE_ENV = "" # TODO: Update this to your prompt environment, ie 'sandbox' | 'latest' | 'production'
# Clients
fp_client = Freeplay(
freeplay_api_key=FREEPLAY_API_KEY, api_base="https://app.freeplay.ai/api"
)
openai_client = OpenAI(api_key=OPENAI_API_KEY)
```
```javascript javascript theme={null}
import "dotenv/config";
import Freeplay from "freeplay";
import OpenAI from "openai";
// TODO: Update these to your environment variables
let FREEPLAY_API_KEY = process.env.FREEPLAY_API_KEY || "";
let OPENAI_API_KEY = process.env.OPENAI_API_KEY || "";
// TODO: Update these to your configuration
let PROJECT_ID = ""; // TODO: Update this to your project ID
let TRACE_DATASET_NAME = ""; // TODO: Update this to your dataset name that targets an agent
let TEST_RUN_NAME = ""; // TODO: Update this to your test name
let TEMPLATE_NAME = ""; // TODO: Update this to your prompt name
let TEMPLATE_ENV = ""; // TODO: Update this to your prompt environment, ie 'sandbox' | 'latest' | 'production'
let API_BASE_URL = "https://app.freeplay.ai/api";
if (!FREEPLAY_API_KEY) throw new Error("FREEPLAY_API_KEY is not set.");
if (!OPENAI_API_KEY) throw new Error("OPENAI_API_KEY is not set.");
// Clients
const fpClient = new Freeplay({
freeplayApiKey: FREEPLAY_API_KEY,
baseUrl: API_BASE_URL,
});
const openaiClient = new OpenAI({ apiKey: OPENAI_API_KEY });
```
```java java theme={null}
package com.freeplay.example;
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.ChatMessage;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.prompts.TemplatePrompt;
import ai.freeplay.client.thin.resources.recordings.CallInfo;
import ai.freeplay.client.thin.resources.recordings.RecordInfo;
import ai.freeplay.client.thin.resources.recordings.ResponseInfo;
import ai.freeplay.client.thin.resources.recordings.TestRunInfo;
import ai.freeplay.client.thin.resources.sessions.Session;
import ai.freeplay.client.thin.resources.sessions.SessionInfo;
import ai.freeplay.client.thin.resources.sessions.TraceInfo;
import ai.freeplay.client.thin.resources.testruns.TestRun;
import ai.freeplay.client.thin.resources.testruns.TraceTestCase;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.core.JsonProcessingException;
import java.net.http.HttpResponse;
import java.util.List;
import java.util.Map;
import java.util.UUID;
import java.util.concurrent.CompletableFuture;
import static ai.freeplay.client.thin.Freeplay.Config;
import static com.freeplay.example.ThinExampleUtils.callAnthropicWithTools;
public class AgentTestRun { // NOTE: This wraps the entire Java example
private static final ObjectMapper objectMapper = new ObjectMapper();
////////////////////////////////////////////////////////
// DOCS EXAMPLE CONFIG
////////////////////////////////////////////////////////
// TODO: Update these to your environment variables
String FREEPLAY_API_KEY = System.getenv("FREEPLAY_API_KEY");
String ANTHROPIC_API_KEY = System.getenv("ANTHROPIC_API_KEY");
TODO: Update these to your configuration
String PROJECT_ID = ""; // TODO: Update this to your project ID
String TRACE_DATASET_NAME = ""; // TODO: Update this to your dataset name that targets an agent
String TEST_RUN_NAME = ""; // TODO: Update this to your test name
String TEMPLATE_NAME = ""; // TODO: Update this to your prompt name
String TEMPLATE_ENV = ""; // TODO: Update this to your prompt environment, ie "sandbox" | "latest" | "production"
String API_BASE = "https://app.freeplay.ai/api";
////////////////////////////////////////////////////////
// Clients
static Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(FREEPLAY_API_KEY)
.baseUrl(API_BASE));
// ... Class continued below
}
```
### Step 2: Minimal Agent Example
For java, the callOpenAIWithTools and callAnthropicWithTools are example classes that can be found [here](https://github.com/freeplayai/freeplay-jvm/blob/main/examples/src/main/java/ai/freeplay/example/java/ThinExampleUtils.java#L66).
```java python theme={null}
def run_agent(
fp_session: SessionInfo,
parent_id: str,
template_name: str,
variables: dict,
test_run_info: Optional[TestRunInfo] = None,
):
# Get prompt from Freeplay
formatted = fp_client.prompts.get_formatted(
project_id=PROJECT_ID,
template_name=template_name,
environment=TEMPLATE_ENV,
variables=variables,
)
model = formatted.prompt_info.model
params = dict(formatted.prompt_info.model_parameters or {})
start = time.time()
completion = openai_client.chat.completions.create(
model=model,
messages=formatted.llm_prompt,
**params,
)
end = time.time()
assistant_msg = completion.choices[0].message
all_messages = formatted.all_messages(assistant_msg)
#################################################
# Handle Agent Activity (ie tool calling, etc.) #
#################################################
# Record to Freeplay
fp_client.recordings.create(
RecordPayload(
project_id=PROJECT_ID,
all_messages=all_messages,
parent_id=parent_id,
inputs=variables,
session_info=fp_session,
test_run_info=test_run_info, # <- NOTE: passing test_run_info links this call to the test run
prompt_version_info=formatted.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted.prompt_info, start_time=start, end_time=end
),
)
)
return assistant_msg.content
```
```javascript javascript theme={null}
async function runAgent(fpSession, parentId, templateName, variables, testRunInfo) {
// Get prompt from Freeplay
const formatted = await fpClient.prompts.getFormatted({
projectId: PROJECT_ID,
templateName,
environment: TEMPLATE_ENV,
variables,
});
const model = formatted.promptInfo.model;
const params = formatted.promptInfo.modelParameters || {};
const startTime = new Date();
const completion = await openaiClient.chat.completions.create({
model,
messages: formatted.llmPrompt,
...params,
});
const endTime = new Date();
const assistantMsg = {
role: completion.choices[0].message.role,
content: completion.choices[0].message.content,
};
const allMessages = formatted.allMessages(assistantMsg);
/*
TODO: Handle Agent Activity (ie tool calling, etc.)
*/
// Record to Freeplay
await fpClient.recordings.create({
projectId: PROJECT_ID,
allMessages,
parentId,
inputs: variables,
sessionInfo: fpSession,
testRunInfo,
promptVersionInfo: formatted.promptInfo,
callInfo: {
provider: formatted.promptInfo.provider,
model: formatted.promptInfo.model,
modelParameters: formatted.promptInfo.modelParameters,
startTime,
endTime,
},
});
return assistantMsg.content;
}
```
```java java-anthropic theme={null}
static String runAgentAnthropic(
SessionInfo sessionInfo,
UUID parentId,
String templateName,
Map variables,
TestRunInfo testRunInfo) throws Exception {
// Get prompt from Freeplay: get() returns CompletableFuture; .get() blocks
TemplatePrompt templatePrompt = fpClient.prompts().get(PROJECT_ID, templateName, TEMPLATE_ENV).get();
FormattedPrompt> formattedPrompt = templatePrompt.bind(new TemplatePrompt.BindRequest(variables)).format();
// Anthropic expects the system content to be passed separately from the messages
String systemContent = formattedPrompt.getSystemContent().orElse(null);
List messages = (List) formattedPrompt.getBoundMessages();
// Call Anthropic API
long start = System.currentTimeMillis();
CompletableFuture> responseFuture = callAnthropicWithTools(
objectMapper,
ANTHROPIC_API_KEY,
formattedPrompt.getPromptInfo().getModel(),
formattedPrompt.getPromptInfo().getModelParameters(),
messages,
systemContent,
formattedPrompt.getToolSchema());
HttpResponse response = responseFuture.get();
long end = System.currentTimeMillis();
// Parse Anthropic response
JsonNode bodyNode;
try {
bodyNode = objectMapper.readTree(response.body());
} catch (JsonProcessingException e) {
throw new RuntimeException("Unable to parse response body.", e);
}
List
### Step 3: Create test run, iterate cases, record outputs
```python python theme={null}
def main():
# Create a Test Run on your dataset (agent/trace)
test_run = fp_client.test_runs.create(
project_id=PROJECT_ID,
testlist=TRACE_DATASET_NAME, # NOTE: the dataset must be created in Freeplay first and have data in it
name=TEST_RUN_NAME,
)
# Iterate test cases
for test_case in tqdm(test_run.trace_test_cases, desc="Running test cases"):
question = getattr(
test_case, "input", ""
) # NOTE: this is the input to the trace
# Create session
session = fp_client.sessions.create()
# Craete the trace
trace: TraceInfo = session.create_trace(
input=question,
agent_name="ExampleAgent",
custom_metadata={"version": "1.0.0"},
)
# NOTE: Prompt variables can be added here if you want to pass them to the prompt
variables = {"user_input": question}
# NOTE: This is the test case ID that will link the recording to the test run
test_run_info = test_run.get_test_run_info(test_case.id)
# Run the agent and log the recording under this test run
assistant_text = run_agent(
fp_session=session,
template_name=TEMPLATE_NAME,
variables=variables,
test_run_info=test_run_info,
parent_id=trace.trace_id,
)
# NOTE: You can attach any evals you compute here
eval_results = {
"evaluation_score": 0.48,
"is_high_quality": True,
}
# NOTE: Record final output for the trace (linked to test run)
trace.record_output(
project_id=PROJECT_ID,
output=assistant_text,
eval_results=eval_results,
test_run_info=test_run_info, # NOTE: passing test_run_info links this call to the test run
)
print("✅ Test run complete. Review results in Freeplay.")
if __name__ == "__main__":
main()
```
```javascript javascript theme={null}
async function main() {
// Create a Test Run on your dataset (agent/trace)
const testRun = await fpClient.testRuns.create({
projectId: PROJECT_ID,
testList: TRACE_DATASET_NAME, // NOTE: the dataset must be created in Freeplay first and have data in it
name: TEST_RUN_NAME,
});
// Iterate test cases
const testCases = testRun.tracesTestCases;
for (let i = 0; i < testCases.length; i++) {
const testCase = testCases[i];
const question = testCase.input || ""; // NOTE: this is the input to the trace
console.log(`Running test case ${i + 1}/${testCases.length}...`);
// Create session
const session = fpClient.sessions.create();
// Create the trace
const trace = session.createTrace({
input: question,
agentName: "ExampleAgent",
customMetadata: { version: "1.0.0" },
});
// NOTE: Prompt variables can be added here if you want to pass them to the prompt
const variables = { user_input: question };
// NOTE: This is the test case ID that will link the recording to the test run
const testRunInfo = {
testRunId: testRun.testRunId,
testCaseId: testCase.id,
};
// Run the agent and log the recording under this test run
const assistantText = await runAgent(
session,
trace.traceId,
TEMPLATE_NAME,
variables,
testRunInfo,
);
// NOTE: You can attach any evals you compute here
const evalResults = {
evaluation_score: 0.48,
is_high_quality: true,
};
// NOTE: Record final output for the trace (linked to test run)
await trace.recordOutput(
PROJECT_ID,
assistantText,
evalResults,
testRunInfo, // NOTE: passing testRunInfo links this call to the test run
);
}
console.log("Test run complete. Review results in Freeplay.");
}
main().catch(console.error);
```
```java java theme={null}
public static void main(String[] args) throws Exception {
if (FREEPLAY_API_KEY == null || FREEPLAY_API_KEY.isEmpty())
throw new RuntimeException("FREEPLAY_API_KEY is not set.");
if (ANTHROPIC_API_KEY == null || ANTHROPIC_API_KEY.isEmpty())
throw new RuntimeException("ANTHROPIC_API_KEY is not set.");
// Create a Test Run on your dataset (agent/trace)
TestRun testRun = fpClient.testRuns().create(
fpClient.testRuns().createRequest(PROJECT_ID, TRACE_DATASET_NAME)
.name(TEST_RUN_NAME)
.build()).get();
// Iterate test cases
List testCases = testRun.getTraceTestCases();
for (int i = 0; i < testCases.size(); i++) {
System.out.printf("Running test case %d/%d...%n", i + 1, testCases.size());
TraceTestCase testCase = testCases.get(i);
String question = testCase.getInput(); // NOTE: this is the input to the trace
// Create session
Session session = fpClient.sessions().create()
.customMetadata(Map.of("customer_id", 123, "is_good", "true"));
// Create the trace
TraceInfo trace = session.createTrace(question)
.agentName("ExampleAgent")
.customMetadata(Map.of("version", "1.0.0"));
// NOTE: Prompt variables can be added here if you want to pass them to the prompt
Map variables = Map.of("question", question);
// NOTE: This is the test case ID that will link the recording to the test run
TestRunInfo testRunInfo = testRun.getTestRunInfo(testCase.getTestCaseId());
// Run the agent and log the recording under this test run
String assistantText = runAgent(
session.getSessionInfo(),
trace.getTraceId(),
TEMPLATE_NAME,
variables,
testRunInfo);
// NOTE: You can attach any evals you compute here
Map evalResults = Map.of(
"evaluation_score", 0.48,
"is_high_quality", true);
// NOTE: Record final output for the trace (linked to test run)
trace.recordOutput(
PROJECT_ID,
assistantText,
evalResults,
testRunInfo // NOTE: passing testRunInfo links this call to the test run
).get();
}
System.out.println("Test run complete. Review results in Freeplay.");
}
// } ...close your class
```
The SDK automatically records all LLM calls, tool invocations, intermediate reasoning steps, and evaluation results throughout the execution.
## Analyzing Results
After running your tests, Freeplay provides comprehensive analysis at both the agent and component levels. The overview shows high-level metrics comparing different versions or models:
You can drill into specific evaluation categories to understand performance across different aspects of your system. Agent evaluations assess the complete workflow:
The row level view reveals how each component contributes to overall system performance. Notice in the example below how we have marked the Claude version as the winner. This view allows us to step through every completion in the dataset and compare side-by-side to see how they differ. You can also note under the session details what actions/steps were taken by the agent in each case during the end-to-end test:
## Best Practices
Include real user interactions that represent typical usage patterns, edge cases that challenge your system, and known failure scenarios that you've encountered. This realistic data ensures your tests catch actual problems users might face.
Run end-to-end tests at critical points in your development cycle. Execute them before deploying to production, after significant code changes, and as part of your CI/CD pipeline. Regular testing catches regressions early when they're easier to fix.
## Advanced Patterns
For multi-agent systems, test the collaboration and handoffs between agents:
```python python theme={null}
for test_case in test_run.trace_test_cases:
# Primary agent processes request
initial_response = primary_agent.process(test_case.input)
# Handoff to specialist if needed
if requires_specialist(initial_response):
final_response = specialist_agent.process(
test_case.input,
context=initial_response
)
```
For RAG pipelines, track each stage of the process:
```python python theme={null}
# Create trace for the complete pipeline
trace_info = session.create_trace(
input=query,
agent_name="rag_pipeline"
)
# Track retrieval, reranking, and generation
retrieved_docs = retrieval_system.search(query)
reranked_docs = reranker.rerank(query, retrieved_docs)
response = generate_response(query, reranked_docs)
trace_info.record_output(
output=response,
eval_results={
'retrieval_relevance': evaluate_retrieval(query, retrieved_docs),
'answer_quality': evaluate_answer(query, response)
}
)
```
***
[Test Runs](/core-concepts/test-runs/test-runs)
[Component Level Test Runs](/core-concepts/test-runs/component-level-test-runs)
# Test Runs
Source: https://docs.freeplay.ai/core-concepts/test-runs/test-runs
Run batch evaluations against datasets to validate performance and catch regressions.
# Test Runs Overview
Test Runs provide structured testing for your AI systems -- aka "evaluations", enabling you to validate performance across datasets, compare versions of your system as you make changes, and catch regressions before they reach production. Freeplay supports two complementary testing approaches designed for different stages of your development workflow, which we call "component" and "end-to-end" tests.
## Core Concepts
A Test Run evaluates your LLM pipeline against a dataset to measure performance, using evaluation metrics you choose and define for your use case. Each run processes your test cases through the pipeline, applies evaluation scores, and provides both aggregate and row-level insights. The foundation of any test is your dataset — a curated collection of scenarios that represent important use cases, edge conditions, and known failure modes.
Test Runs integrate with Freeplay's overall evaluation system, allowing you to apply model-graded evals, code-based checks, and human review to assess quality. You can compare results across different versions of your prompts, tools, models, or other parts of your system by creating separate Test Runs for each version you want to compare, then comparing scores head to head (either in the Freeplay UI, or your code).
## When to Use Each Approach
The choice between component and end-to-end testing depends on what you're trying to validate. Our suggestinon:
* Use component testing when iterating on a specific prompt, model, or tool to validate your changes in isolation
* Use end-to-end testing for your entire code path to test how a given set of inputs turns into outputs for your agent or system
Together these create a helpful two-step workflow for many changes: Once your component changes are working well, then use end-to-end testing to validate how those changes behave within your complete system. End-to-end tests are essential before deploying to production, making system architecture changes, or validating agent behavior with tool usage.
## Testing Approaches
### Component-Level Test Runs
Component testing focuses on individual prompts and model changes in isolation. This approach enables rapid iteration during prompt engineering without the overhead of running your entire system. You can execute component tests through either the Freeplay UI for quick no-code testing, or the SDK for programmatic control.
Component tests excel at prompt optimization, model comparisons, parameter tuning, and enabling non-technical team members to participate in the iteration process.
[Learn more about Component Testing →](/core-concepts/test-runs/component-level-test-runs)
### End-to-End Test Runs
End-to-end testing validates your complete system workflow including agents, RAG pipelines, and multi-step processes. When you change any component —- whether it's a prompt, tool, or orchestration logic -— you need to understand how it affects your final output. These tests execute through the SDK in order to test your actual system.
End-to-end tests are essential for production validation, agent systems with multiple prompts and tools, RAG pipelines, and multi-turn conversations. They capture the full complexity of your system as users experience it. Consider logging a specific commit name or git SHA with these tests to connect them to your code.
[Learn more about End-to-End Testing →](/core-concepts/test-runs/end-to-end-test-runs)
## Getting Started
For developers, begin by installing the Freeplay SDK and creating datasets from your production data. Set up both end-to-end and component tests as part of your development workflow, and integrat them into your CI/CD pipeline for automated validation when ready.
Product teams can jump straight into the Freeplay UI to create test datasets from important use cases and run component tests on prompt changes. The visual interface makes it easy to review results and provide feedback without technical expertise.
**API Reference**: See the [Create Test Run](/api-reference/test-runs/create-test-run), [List Test Runs](/api-reference/test-runs/list-test-runs), and [Get Test Run Results](/api-reference/test-runs/get-test-run-results) endpoints. Test runs work with both [Prompt Datasets](/api-reference/prompt-datasets/get-prompt-dataset) and [Agent Datasets](/api-reference/agent-datasets/list-agent-datasets).
***
What's Next
Now that you're armed with the ability to test your models, let's move onto Datasets.
* [Component Level Test Runs](/core-concepts/test-runs/component-level-test-runs)
* [End-to-End Test Runs](/core-concepts/test-runs/end-to-end-test-runs)
* [Datasets](/core-concepts/datasets/datasets)
* [Evaluations](/core-concepts/evaluations/evaluations)
# API Reference
Source: https://docs.freeplay.ai/developer-resources/api-reference
Complete reference for the Freeplay HTTP API including authentication, endpoints, and examples
We're transitioning to OpenAPI spec-driven documentation. The new [API Reference](/openapi/introduction) features interactive "Try It" functionality and auto-generated examples, and you can access the full OpenAPI spec there. This page will remain available for detailed descriptions during the transition.
***
The Freeplay HTTP API provides programmatic access to all platform capabilities. While the [Freeplay SDKs](/freeplay-sdk/organizing-principles) offer language-native bindings for common operations, the API exposes the full range of functionality.
**SDK vs. API**: The Freeplay SDKs are designed to cover the most common integration patterns—prompt management, recording completions, and running tests. The HTTP API is a superset that includes additional capabilities like bulk operations, advanced search, and administrative endpoints.
## Getting Started
### Base URL
Your API root is your Freeplay instance URL plus `/api/v2/`.
| Deployment Type | Base URL |
| --------------- | --------------------------------------------- |
| Cloud | `https://app.freeplay.ai/api/v2` |
| Private | `https://{your-subdomain}.freeplay.ai/api/v2` |
### Authentication
Authenticate requests using your API key in the Authorization header:
```bash theme={null}
Authorization: Bearer {freeplay_api_key}
```
API keys are managed at `https://app.freeplay.ai/settings/api-access`.
### Error Handling
Freeplay uses standard HTTP status codes:
| Code | Description |
| ----- | -------------------------------------------------- |
| `200` | Success |
| `400` | Bad Request - Malformed request or invalid data |
| `401` | Unauthorized - Invalid or missing API key |
| `404` | Not Found - Resource doesn't exist or no access |
| `500` | Server Error - Transient issue, retry with backoff |
For 500 errors, retry up to three times with at least 5 seconds between attempts. Do not retry 400-level errors—these indicate client issues that won't resolve on retry.
**Example Error Response:**
```json theme={null}
{
"message": "Session ID 456 is not a valid uuid4 format."
}
```
***
## Observability
Freeplay organizes observability data in a hierarchy:
```
Session (user conversation or workflow)
└── Trace (optional grouping of related completions)
└── Completion (a single LLM call)
```
* **Sessions** are created implicitly when you record your first completion with a session ID, or you can create one
* **Traces** are optional but required for agent workflows—they group related completions together
* **Completions** are the atomic unit: a prompt sent to an LLM and its response
For conceptual background, see [Sessions, Traces, and Completions](/core-concepts/observability/sessions-traces-and-completions). For SDK usage, see [Organizing Principles](/freeplay-sdk/organizing-principles).
### Sessions
Sessions group related LLM completions together—typically representing a user conversation or workflow.
**Base URL**: `/api/v2/projects//sessions`
Sessions are created implicitly when you record a completion. You only need to generate a session ID (UUID v4) client-side. For SDK usage, see [Sessions](/freeplay-sdk/sessions).
#### Retrieve Sessions
`GET /`
Returns sessions with their completions, ordered by most recent first.
**Query Parameters:**
| Parameter | Type | Description | Required |
| ------------------- | ------ | -------------------------------------------- | ------------------ |
| `page` | `int` | Page number | No (Default: `1`) |
| `page_size` | `int` | Results per page (max: 100) | No (Default: `10`) |
| `from_date` | `str` | Start date (inclusive). Format: `YYYY-MM-DD` | No |
| `to_date` | `str` | End date (exclusive). Format: `YYYY-MM-DD` | No |
| `test_list` | `str` | Filter by dataset name | No |
| `test_run_id` | `str` | Filter by test run ID | No |
| `prompt_name` | `str` | Filter by prompt template name | No |
| `review_queue_id` | `str` | Filter by review queue ID | No |
| `custom_metadata.*` | `dict` | Filter by session metadata | No |
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/sessions?page_size=20&from_date=2025-01-01"
```
**Response:**
```json theme={null}
[
{
"session_id": "85d7d393-4e85-4664-8598-5dd91dc75b5b",
"start_time": "2024-07-05T14:33:02.721000",
"custom_metadata": {},
"messages": [
{
"completion_id": "0202ae57-1098-4d3e-94f7-1fbb034fbed5",
"prompt_template_name": "my-prompt",
"model_name": "claude-2.1",
"provider_name": "anthropic",
"environment": "latest",
"prompt": [...],
"response": "You bundle prompts by...",
"input_variables": {"question": "How do I bundle prompts?"}
}
]
}
]
```
#### Delete Session
`DELETE /`
Permanently deletes a session and all associated completions.
This operation cannot be undone.
```bash theme={null}
curl -X DELETE \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/sessions/$SESSION_ID"
```
### Traces
Traces group related completions within a session. They're essential for agent workflows where multiple LLM calls work together to accomplish a task.
**Base URL**: `/api/v2/projects//sessions//traces`
For SDK usage, see [Traces](/freeplay-sdk/traces).
#### Record a Trace
`POST /id/`
Records a trace within a session. Like sessions, traces are created implicitly—generate a UUID v4 client-side and use it as the trace ID.
**Request Payload:**
| Parameter | Type | Description | Required |
| ----------------- | ---------------------------------- | ------------------------------------------------ | -------- |
| `input` | `any` | Input to the trace (e.g., user query) | Yes |
| `output` | `any` | Output from the trace (e.g., final response) | Yes |
| `agent_name` | `str` | Name of the agent (required for agent workflows) | No |
| `name` | `str` | Display name for the trace | No |
| `custom_metadata` | `dict[str, str\|int\|float\|bool]` | Custom metadata | No |
| `eval_results` | `dict[str, float\|bool]` | Code evaluation results | No |
| `parent_id` | `UUID` | Parent trace ID for nested traces | No |
| `test_run_info` | `{test_run_id, test_case_id}` | Test run association | No |
| `kind` | `"tool"` | Set to `"tool"` for tool call traces | No |
| `start_time` | `datetime` | Trace start time | No |
| `end_time` | `datetime` | Trace end time | No |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/sessions/$SESSION_ID/traces/id/$TRACE_ID" \
-d '{
"input": "What is the weather in San Francisco?",
"output": "The weather in San Francisco is 65°F and sunny.",
"agent_name": "weather-agent",
"custom_metadata": {"version": "1.0.0"}
}'
```
When building agents, record completions with `trace_info` to associate them with a trace, then call this endpoint to finalize the trace with its output.
### Completions
Completions are the atomic unit of observability—a single LLM call with its prompt and response.
**Base URL**: `/api/v2/projects//sessions//completions`
For SDK usage, see [Recording Completions](/freeplay-sdk/recording-completions).
#### Record a Completion
`POST /`
Records an LLM completion to a session. This is the primary endpoint for logging LLM interactions.
**Request Payload:**
| Parameter | Type | Description | Required |
| --------------- | ----------------------------------------------------- | ------------------------------------------ | -------- |
| `messages` | `list[{role, content}]` | Messages sent to and received from the LLM | Yes |
| `inputs` | `dict[str, any]` | Input variables used in the prompt | Yes |
| `prompt_info` | `{prompt_template_version_id, environment}` | Prompt template info | Yes |
| `trace_info` | `{trace_id}` | Associate completion with a trace | No |
| `tool_schema` | `list[{name, description, parameters}]` | Tool definitions | No |
| `session_info` | `{custom_metadata: dict}` | Session metadata | No |
| `call_info` | `{start_time, end_time, model, provider, usage, ...}` | LLM call details | No |
| `test_run_info` | `{test_run_id, test_case_id}` | Test run association | No |
| `completion_id` | `UUID` | Custom completion ID | No |
| `eval_results` | `dict[str, bool\|float]` | Code evaluation results | No |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/sessions/$SESSION_ID/completions" \
-d '{
"messages": [
{"role": "user", "content": "Generate an album name for Taylor Swift"},
{"role": "assistant", "content": "Rainy Melodies"}
],
"inputs": {"pop_star": "Taylor Swift"},
"prompt_info": {
"prompt_template_version_id": "f503c15e-2f0f-4ce4-b443-4c87d0b6435d",
"environment": "prod"
}
}'
```
**Response:**
```json theme={null}
{
"completion_id": "707bc301-85e9-4f02-aa97-faba8cd7774a"
}
```
**With Trace Association:**
To associate a completion with a trace (for agent workflows), include `trace_info`:
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/sessions/$SESSION_ID/completions" \
-d '{
"messages": [...],
"inputs": {...},
"prompt_info": {...},
"trace_info": {
"trace_id": "abc123-trace-uuid"
}
}'
```
***
## Prompt Templates
Create, retrieve, and manage prompt templates programmatically. For conceptual background on prompt management patterns, see [Prompt Management](/core-concepts/prompt-management/managing-prompts).
**Base URL**: `/api/v2/projects//prompt-templates`
### Create Prompt Template
`POST /`
Creates a new prompt template (without any versions). Typically used when you want to create a template first, then add versions separately.
NOTE: The same objective can be accomplished using the `create_template_if_not_exists` parameter on the [Create Version by Name](#create-version-by-name) endpoint below.
**Request Payload:**
| Parameter | Type | Description | Required |
| --------- | ------ | ------------------------------------- | -------- |
| `name` | `str` | Template name (unique within project) | Yes |
| `id` | `uuid` | Custom template ID | No |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates" \
-d '{"name": "my-assistant"}'
```
**Response:**
```json theme={null}
{
"id": "5cccc8e6-b163-4094-8bd2-90030f151ec8"
}
```
### List Prompt Templates
`GET /`
Returns all prompt templates in a project with pagination.
**Query Parameters:**
| Parameter | Type | Description | Default |
| ----------- | ----- | --------------------------- | ------- |
| `page` | `int` | Page number | `1` |
| `page_size` | `int` | Results per page (max: 100) | `30` |
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates?page=1&page_size=50"
```
### Create Version by Name
`POST /name//versions`
Creates a new prompt version by template name. **This is the recommended endpoint for CI/CD workflows** because it supports creating the template automatically if it doesn't exist.
**Query Parameters:**
| Parameter | Type | Description | Default |
| ------------------------------- | --------- | ---------------------------- | ------- |
| `create_template_if_not_exists` | `boolean` | Create template if not found | `false` |
**Request Payload:**
| Parameter | Type | Description | Required |
| --------------------- | ----------------------- | ------------------------------------------ | -------- |
| `template_messages` | `list[{role, content}]` | Message array with mustache variables | Yes |
| `model` | `str` | Model name | Yes |
| `provider` | `str` | Provider key (`openai`, `anthropic`, etc.) | Yes |
| `llm_parameters` | `dict` | Model parameters (`temperature`, etc.) | No |
| `tool_schema` | `list` | Tool definitions | No |
| `output_schema` | `dict` | Structured output schema | No |
| `version_name` | `str` | Display name | No |
| `version_description` | `str` | Description | No |
| `environments` | `list[str]` | Environments to deploy to | No |
```bash cURL theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/name/my-assistant/versions?create_template_if_not_exists=true" \
-d '{
"template_messages": [
{"role": "system", "content": "You are a helpful assistant. The user'\''s name is {{user_name}}."},
{"role": "user", "content": "{{user_input}}"}
],
"provider": "openai",
"model": "gpt-4o",
"llm_parameters": {"temperature": 0.2, "max_tokens": 1024},
"version_name": "v1.2.0",
"version_description": "Production release with improved system prompt"
}'
```
```python Python theme={null}
import os
import requests
FREEPLAY_API_KEY = os.getenv("FREEPLAY_API_KEY")
project_id = os.getenv("FREEPLAY_PROJECT_ID")
base_url = "https://app.freeplay.ai/api/v2"
def sync_prompt(template_name: str, messages: list, model: str, provider: str = "openai"):
"""Create or update a prompt template version."""
url = f"{base_url}/projects/{project_id}/prompt-templates/name/{template_name}/versions"
response = requests.post(
url,
params={"create_template_if_not_exists": "true"},
headers={
"Authorization": f"Bearer {FREEPLAY_API_KEY}",
"Content-Type": "application/json"
},
json={
"template_messages": messages,
"provider": provider,
"model": model,
"llm_parameters": {"temperature": 0.2, "max_tokens": 1024},
"version_name": "v1.0.0"
}
)
response.raise_for_status()
return response.json()
# Example usage
result = sync_prompt(
template_name="my-assistant",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "{{user_input}}"}
],
model="gpt-4o"
)
print(f"Created version: {result['prompt_template_version_id']}")
```
```typescript TypeScript theme={null}
const FREEPLAY_API_KEY = process.env.FREEPLAY_API_KEY;
const projectId = process.env.FREEPLAY_PROJECT_ID;
const baseUrl = "https://app.freeplay.ai/api/v2";
async function syncPrompt(
templateName: string,
messages: Array<{role: string; content: string}>,
model: string,
provider: string = "openai"
) {
const url = `${baseUrl}/projects/${projectId}/prompt-templates/name/${templateName}/versions?create_template_if_not_exists=true`;
const response = await fetch(url, {
method: "POST",
headers: {
Authorization: `Bearer ${FREEPLAY_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
template_messages: messages,
provider,
model,
llm_parameters: { temperature: 0.2, max_tokens: 1024 },
version_name: "v1.0.0"
})
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
return response.json();
}
// Example usage
const result = await syncPrompt(
"my-assistant",
[
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "{{user_input}}" }
],
"gpt-4o"
);
console.log(`Created version: ${result.prompt_template_version_id}`);
```
**Response:**
```json theme={null}
{
"prompt_template_id": "5cccc8e6-b163-4094-8bd2-90030f151ec8",
"prompt_template_version_id": "f503c15e-2f0f-4ce4-b443-4c87d0b6435d",
"prompt_template_name": "my-assistant",
"version_name": "v1.2.0",
"version_description": "Production release with improved system prompt",
"format_version": 2,
"project_id": "abc123-project-uuid",
"content": [...],
"metadata": {
"flavor": "openai_chat",
"model": "gpt-4o",
"provider": "openai",
"params": {"temperature": 0.2, "max_tokens": 1024}
}
}
```
**Code-managed prompts**: Use the `create_template_if_not_exists=true` parameter in your CI/CD pipeline to automatically sync prompts from your codebase to Freeplay. See [Code as source of truth](/core-concepts/prompt-management/managing-prompts#code-as-the-source-of-truth-for-prompts) for the full workflow.
### Retrieve by Name
`POST /name/`
Fetches a prompt template in one of three forms:
| Form | Description | How to Request |
| ------------- | -------------------------------------------- | ------------------------------ |
| **Raw** | Template with `{{variable}}` placeholders | No body, no `format` param |
| **Bound** | Variables inserted, provider-agnostic format | Pass variables in body |
| **Formatted** | Variables inserted, provider-specific format | Pass variables + `format=true` |
**Query Parameters:**
| Parameter | Type | Description | Default |
| ------------- | --------- | ----------------------- | -------- |
| `environment` | `str` | Environment tag | `latest` |
| `format` | `boolean` | Return formatted prompt | `false` |
**Formatted Prompt (most common):**
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/name/album_bot?environment=prod&format=true" \
-d '{"pop_star": "Taylor Swift"}'
```
**Response:**
```json theme={null}
{
"format_version": 2,
"prompt_template_id": "5cccc8e6-b163-4094-8bd2-90030f151ec8",
"prompt_template_name": "album_bot",
"prompt_template_version_id": "f503c15e-2f0f-4ce4-b443-4c87d0b6435d",
"formatted_content": [
{
"role": "user",
"content": "Generate a two word album name in the style of Taylor Swift"
}
],
"formatted_tool_schema": [...],
"metadata": {
"flavor": "openai_chat",
"model": "gpt-3.5-turbo-0125",
"provider": "openai",
"params": {"max_tokens": 100, "temperature": 0.2}
}
}
```
### Retrieve by Version ID
`POST /id//versions/`
Fetch a specific prompt version. Useful for pinning to a known version.
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/id/$TEMPLATE_ID/versions/$VERSION_ID"
```
### Retrieve All Templates
`GET /all/`
Returns all prompt templates in an environment.
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/all/prod"
```
### Create Version by ID
`POST /id//versions`
Adds a new version to an existing prompt template (identified only by ID). Deployed to `latest` by default.
**Request Payload:**
| Parameter | Type | Description | Required |
| --------------------- | ----------------------- | ------------------------------------------ | -------- |
| `template_messages` | `list[{role, content}]` | Message array with mustache variables | Yes |
| `model` | `str` | Model name | Yes |
| `provider` | `str` | Provider key (`openai`, `anthropic`, etc.) | Yes |
| `llm_parameters` | `dict` | Model parameters (`temperature`, etc.) | No |
| `tool_schema` | `list` | Tool definitions | No |
| `output_schema` | `dict` | Structured output schema | No |
| `version_name` | `str` | Display name | No |
| `version_description` | `str` | Description | No |
| `environments` | `list[str]` | Environments to deploy to | No |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/id/$TEMPLATE_ID/versions" \
-d '{
"template_messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "{{user_input}}"}
],
"provider": "openai",
"model": "gpt-4o",
"llm_parameters": {"temperature": 0.2, "max_tokens": 256}
}'
```
### Update Environments
`POST /id//versions//environments`
Assign a version to additional environments.
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/id/$TEMPLATE_ID/versions/$VERSION_ID/environments" \
-d '{"environments": ["staging", "prod"]}'
```
For SDK-based prompt management, see [Prompts](/freeplay-sdk/prompts). For the conceptual guide on managing prompts programmatically, see [Code as source of truth](/core-concepts/prompt-management/managing-prompts#code-as-the-source-of-truth-for-prompts).
***
## Test Runs
Execute batch tests using saved datasets.
**Base URL**: `/api/v2/projects//test-runs`
### Create Test Run
`POST /`
Creates a new test run from an existing dataset.
| Parameter | Type | Description | Required |
| ---------------------- | --------- | ------------------------ | -------------------- |
| `dataset_name` | `str` | Name of the dataset | Yes |
| `include_outputs` | `boolean` | Include expected outputs | No (Default: `true`) |
| `test_run_name` | `str` | Display name | No |
| `test_run_description` | `str` | Description | No |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/test-runs" \
-d '{"dataset_name": "Example Tests"}'
```
**Response:**
```json theme={null}
{
"test_run_id": "bd3eb06c-f93b-46a4-aa3b-d240789c8a06",
"test_run_name": "",
"test_run_description": "",
"test_cases": [
{
"test_case_id": "91e60c9e-fbaa-4990-b4cc-7a8bd067f298",
"variables": {"question": "How do test runs work?"},
"output": null
}
]
}
```
### Retrieve Test Run Results
`GET /id/`
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/test-runs/id/$TEST_RUN_ID"
```
**Response:**
```json theme={null}
{
"id": "2a9dd8bd-6c29-47c8-9ca4-427f73174881",
"name": "regression-test",
"model_name": "gpt-4o-mini",
"prompt_name": "rag-qa",
"sessions_count": 132,
"summary_statistics": {
"auto_evaluation": {"Answer Accuracy": {"5": 104, "4": 9}},
"client_evaluation": {},
"human_evaluation": {}
}
}
```
### List Test Runs
`GET /`
| Parameter | Type | Description | Default |
| ----------- | ----- | --------------------------- | ------- |
| `page` | `int` | Page number | `1` |
| `page_size` | `int` | Results per page (max: 100) | `100` |
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/test-runs"
```
For SDK-based testing, see [Test Runs](/freeplay-sdk/test-runs).
***
## Customer Feedback
Record user feedback for completions and traces.
### Completion Feedback
**Base URL**: `/api/v2/projects//completion-feedback`
`POST /id/`
| Parameter | Type | Description | Required |
| ------------------- | ----------------------- | ---------------------------- | -------- |
| `freeplay_feedback` | `str` | `"positive"` or `"negative"` | Yes |
| `*` | `str\|float\|int\|bool` | Custom feedback attributes | No |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/completion-feedback/id/$COMPLETION_ID" \
-d '{
"freeplay_feedback": "positive",
"rating": 5,
"comment": "Great response"
}'
```
### Trace Feedback
**Base URL**: `/api/v2/projects//trace-feedback`
`POST /id/`
Same parameters as completion feedback.
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/trace-feedback/id/$TRACE_ID" \
-d '{"freeplay_feedback": "negative", "reason": "incomplete answer"}'
```
For SDK-based feedback, see [Customer Feedback](/freeplay-sdk/customer-feedback).
***
## Search API
Query sessions, traces, and completions with powerful filtering. This functionality is API-only and not available through the SDKs.
### Endpoints
| Endpoint | Description |
| -------------------------- | ------------------ |
| `POST /search/sessions` | Search sessions |
| `POST /search/traces` | Search traces |
| `POST /search/completions` | Search completions |
All endpoints support pagination via `page` and `page_size` query parameters.
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/search/completions?page=1&page_size=20" \
-d '{"filters": {"field": "cost", "op": "gte", "value": 0.01}}'
```
### Filter Operators
| Operator | Description |
| ---------- | --------------------- |
| `eq` | Equals |
| `lt` | Less than |
| `gt` | Greater than |
| `lte` | Less than or equal |
| `gte` | Greater than or equal |
| `contains` | Contains substring |
| `between` | Within numeric range |
### Available Filters
| Field | Supported Operators | Example Value |
| ---------------------------------------- | ------------------------------------------ | --------------------------- |
| `cost` | `eq`, `lt`, `gt`, `lte`, `gte` | `0.003` |
| `latency` | `eq`, `lt`, `gt`, `lte`, `gte` | `8` |
| `start_time` | `eq`, `lt`, `gt`, `lte`, `gte` | `"2024-06-01 00:00:00"` |
| `environment` | `eq` | `"staging"` |
| `prompt_template` | `eq` | `"my-prompt"` |
| `prompt_template_id` | `eq` | `"uuid..."` |
| `model` | `eq` | `"gpt-4o"` |
| `provider` | `eq` | `"openai"` |
| `review_status` | `eq` | `"review_complete"` |
| `agent_name` | `eq` | `"support-agent"` |
| `trace_agent_name` | `eq` | `"my-agent"` |
| `api_key` | `eq` | `"production-key"` |
| `assignee` | `eq` | `"user@example.com"` |
| `review_theme` | `eq` | `"Response Quality Issues"` |
| `completion_output` | `contains` | `"weather"` |
| `completion_inputs.*` | `contains` | `"topic": "weather"` |
| `completion_feedback.*` | `contains` | `"rating": "positive"` |
| `session_custom_metadata.*` | `contains` | `"user_type": "premium"` |
| `trace_custom_metadata.*` | `contains` | `"workflow": "onboarding"` |
| `trace_input.*` | `contains` | `"query": "weather"` |
| `trace_output.*` | `contains` | `"response": "sunny"` |
| `trace_feedback.*` | `contains` | `"rating": "positive"` |
| `completion_evaluation_results.*` | `eq` | `"Response Quality": "4"` |
| `completion_client_evaluation_results.*` | `eq` | `"score": "85"` |
| `trace_evaluation_results.*` | `eq`, `gt`, `lt`, `gte`, `lte`, `contains` | `"Quality Score": 5` |
| `trace_client_eval_results.*` | `eq`, `contains` | `"confidence_score": 0.95` |
| `evaluation_notes.content` | `contains` | `"needs review"` |
| `evaluation_notes.author` | `eq` | `"user@example.com"` |
| `evaluation_notes.created_at` | `gt`, `lt`, `gte`, `lte` | `"2024-06-01 00:00:00"` |
### Compound Filters
Combine filters using `and`, `or`, and `not`:
```json theme={null}
{
"filters": {
"and": [
{"field": "cost", "op": "gte", "value": 0.001},
{
"or": [
{"field": "model", "op": "eq", "value": "gpt-4o"},
{"field": "model", "op": "eq", "value": "claude-3-opus"}
]
},
{
"not": {"field": "environment", "op": "eq", "value": "prod"}
}
]
}
}
```
***
## Additional API Endpoints
The following endpoints provide administrative and bulk operations not covered by the SDKs.
### Projects
**Base URL**: `/api/v2/projects`
#### List All Projects
`GET /all`
Returns all projects in your workspace.
Does not work with project-scoped API keys. Private projects are excluded.
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/all"
```
### Agents
**Base URL**: `/api/v2/projects//agents`
#### List Agents
`GET /`
| Parameter | Type | Description | Default |
| ----------- | -------- | --------------------------- | ------- |
| `page` | `int` | Page number | `1` |
| `page_size` | `int` | Results per page (max: 100) | `30` |
| `name` | `string` | Filter by exact name | - |
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/agents"
```
### Datasets
**Base URL**: `/api/v2/projects//datasets`
#### Retrieve Dataset Metadata
`GET /name/` or `GET /id/`
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/datasets/name/Sample"
```
#### Retrieve Dataset Test Cases
`GET /name//test-cases` or `GET /id//test-cases`
```bash theme={null}
curl -H "Authorization: Bearer $FREEPLAY_API_KEY" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/datasets/name/Sample/test-cases"
```
#### Upload Test Cases
`POST /id//test-cases`
Maximum 100 test cases per request.
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/datasets/id/$DATASET_ID/test-cases" \
-d '{
"examples": [
{"inputs": {"question": "What is Freeplay?"}, "output": "An LLM platform"},
{"inputs": {"question": "How do I integrate?"}, "output": "Use the SDK"}
]
}'
```
### Completions Statistics
**Base URL**: `/api/v2/projects//completions`
#### Aggregate Statistics
`POST /statistics`
Returns evaluation statistics across all prompts for a date range (max 30 days).
| Parameter | Type | Description | Default |
| ----------- | ----- | ---------------------- | ---------- |
| `from_date` | `str` | Start date (inclusive) | 7 days ago |
| `to_date` | `str` | End date (exclusive) | Today |
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/completions/statistics" \
-d '{"from_date": "2025-01-01", "to_date": "2025-01-15"}'
```
#### Statistics by Prompt
`POST /statistics/`
Same parameters as aggregate statistics, filtered to a specific prompt template.
***
## Complete Examples
### End-to-End LLM Interaction
This example fetches a prompt, calls OpenAI, and records the completion:
```python theme={null}
import requests
import os
import json
import uuid
project_id = os.getenv("FREEPLAY_PROJECT_ID")
api_root = f"https://app.freeplay.ai/api/v2/projects/{project_id}"
headers = {"Authorization": f"Bearer {os.getenv('FREEPLAY_API_KEY')}"}
# 1. Fetch formatted prompt
prompt_resp = requests.post(
f"{api_root}/prompt-templates/name/album_bot",
headers=headers,
params={"environment": "prod", "format": "true"},
json={"pop_star": "Taylor Swift"}
)
formatted_prompt = prompt_resp.json()
# 2. Call OpenAI
openai_resp = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={
"Authorization": f"Bearer {os.getenv('OPENAI_API_KEY')}",
"Content-Type": "application/json"
},
json={
"model": formatted_prompt["metadata"]["model"],
"messages": formatted_prompt["formatted_content"],
**formatted_prompt["metadata"]["params"]
}
)
response_message = openai_resp.json()['choices'][0]['message']
# 3. Record to Freeplay
messages = formatted_prompt["formatted_content"] + [response_message]
session_id = str(uuid.uuid4())
requests.post(
f"{api_root}/sessions/{session_id}/completions",
headers={**headers, "Content-Type": "application/json"},
json={
"messages": messages,
"inputs": {"pop_star": "Taylor Swift"},
"prompt_info": {
"prompt_template_version_id": formatted_prompt["prompt_template_version_id"],
"environment": "prod"
}
}
)
```
### Executing a Test Run
```python theme={null}
import requests
import os
import json
import uuid
project_id = os.getenv("FREEPLAY_PROJECT_ID")
api_root = f"https://app.freeplay.ai/api/v2/projects/{project_id}"
headers = {"Authorization": f"Bearer {os.getenv('FREEPLAY_API_KEY')}"}
# Create test run
test_run_resp = requests.post(
f"{api_root}/test-runs",
headers={**headers, "Content-Type": "application/json"},
json={"dataset_name": "Example Tests"}
)
test_run = test_run_resp.json()
# Process each test case
for test_case in test_run["test_cases"]:
# Fetch prompt with test case variables
prompt_resp = requests.post(
f"{api_root}/prompt-templates/name/rag-qa",
headers=headers,
params={"environment": "prod", "format": "true"},
json=test_case['variables']
)
formatted_prompt = prompt_resp.json()
# Call LLM (example with OpenAI)
openai_resp = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={
"Authorization": f"Bearer {os.getenv('OPENAI_API_KEY')}",
"Content-Type": "application/json"
},
json={
"model": formatted_prompt["metadata"]["model"],
"messages": formatted_prompt["formatted_content"],
**formatted_prompt["metadata"]["params"]
}
)
response_message = openai_resp.json()['choices'][0]['message']
# Record with test run info
messages = formatted_prompt["formatted_content"] + [response_message]
session_id = str(uuid.uuid4())
requests.post(
f"{api_root}/sessions/{session_id}/completions",
headers={**headers, "Content-Type": "application/json"},
json={
"messages": messages,
"inputs": test_case['variables'],
"prompt_info": {
"prompt_template_version_id": formatted_prompt["prompt_template_version_id"],
"environment": "prod"
},
"test_run_info": {
"test_run_id": test_run["test_run_id"],
"test_case_id": test_case["test_case_id"]
}
}
)
```
### Agent Workflow with Traces
This example shows how to record an agent workflow with multiple completions grouped by a trace:
```python theme={null}
import requests
import os
import uuid
project_id = os.getenv("FREEPLAY_PROJECT_ID")
api_root = f"https://app.freeplay.ai/api/v2/projects/{project_id}"
headers = {"Authorization": f"Bearer {os.getenv('FREEPLAY_API_KEY')}"}
# Generate IDs for session and trace
session_id = str(uuid.uuid4())
trace_id = str(uuid.uuid4())
user_query = "What's the weather in San Francisco and should I bring an umbrella?"
# 1. First LLM call - agent decides to check weather
prompt_resp = requests.post(
f"{api_root}/prompt-templates/name/weather-agent",
headers=headers,
params={"environment": "prod", "format": "true"},
json={"query": user_query}
)
formatted_prompt = prompt_resp.json()
# Call LLM...
# response_1 = call_llm(formatted_prompt)
# Record first completion with trace association
requests.post(
f"{api_root}/sessions/{session_id}/completions",
headers={**headers, "Content-Type": "application/json"},
json={
"messages": [...], # Include prompt + response
"inputs": {"query": user_query},
"prompt_info": {
"prompt_template_version_id": formatted_prompt["prompt_template_version_id"],
"environment": "prod"
},
"trace_info": {"trace_id": trace_id} # Associate with trace
}
)
# 2. Second LLM call - agent formats final response
# ... make another LLM call and record with same trace_id ...
# 3. Finalize the trace with input/output
requests.post(
f"{api_root}/sessions/{session_id}/traces/id/{trace_id}",
headers={**headers, "Content-Type": "application/json"},
json={
"input": user_query,
"output": "The weather in San Francisco is 65°F and sunny. No umbrella needed!",
"agent_name": "weather-agent"
}
)
```
***
## Next Steps
* [SDK Setup](/freeplay-sdk/setup) - Language-native SDK installation
* [Recording Completions](/freeplay-sdk/recording-completions) - SDK-based observability
* [Traces](/freeplay-sdk/traces) - Grouping completions for agent workflows
* [Prompt Management](/core-concepts/prompt-management/managing-prompts) - Managing prompts in Freeplay
* [Test Runs](/core-concepts/test-runs/test-runs) - Batch testing concepts
# Google Agent Development Kit (ADK)
Source: https://docs.freeplay.ai/developer-resources/integrations/adk
Integrate Google's Agent Development Kit with Freeplay for agent observability and evaluation.
# Google Agent Observability and Evaluation with Freeplay
See Google's Guide [here](https://google.github.io/adk-docs/observability/freeplay/).
We provide an end-to-end integrated workflow for building and optimizing AI agents with Google's ADK. With ADK and Freeplay your whole team can easily collaborate to iterate on agent instructions (prompts), experiment with and compare different models and agent changes, run evals both offline and online to measure quality, monitor production, and review data by hand.
Key benefits of connecting Freeplay:
* **Simple observability** - focused on agents, LLM calls and tool calls for easy human review
* **Online evals/automated scorers** - for error detection in production
* **Offline evals and experiment comparison** - to test changes before deploying
* **Prompt management** - supports pushing changes straight from the Freeplay playground to code
* **Human review workflow** - for collaboration on error analysis and data annotation
* **Powerful UI** - makes it possible for domain experts to collaborate closely with engineers
Freeplay and Google's ADK complement one another. ADK gives you a powerful and expressive agent orchestration framework while Freeplay plugs in for observability, prompt management, evaluation and testing. Once you integrate with Freeplay, you can update prompts and evals from the Freeplay UI or from code, so that anyone on your team can contribute.
## Getting Started
Below is a guide for getting started with Freeplay and ADK. You can also find a full sample ADK agent repo [here](https://github.com/228Labs/freeplay-google-demo).
### Create a Freeplay Account
Sign up for a free [Freeplay account](https://freeplay.ai/signup) .
After creating an account, you can define the following environment variables:
```
FREEPLAY_PROJECT_ID=
FREEPLAY_API_KEY=
FREEPLAY_API_URL=
```
### Use Freeplay ADK Library
Install the Freeplay ADK library:
```bash theme={null}
pip install freeplay-python-adk
```
Freeplay will automatically capture OTel logs from your ADK application when you initialize observability:
```python python theme={null}
from freeplay_python_adk.client import FreeplayADK
FreeplayADK.initialize_observability()
```
You'll also want to pass in the Freeplay plugin to your App:
```python python theme={null}
from app.agent import root_agent
from freeplay_python_adk.freeplay_observability_plugin import FreeplayObservabilityPlugin
from google.adk.runners import App
app = App(
name="app",
root_agent=root_agent,
plugins=[FreeplayObservabilityPlugin()],
)
__all__ = ["app"]
```
You can now use ADK as you normally would, and you will see logs flowing to Freeplay in the Observability section.
## Observability
Freeplay's Observability feature gives you a clear view into how your agent is behaving in production. You can dig into to individual agent traces to understand each step and diagnose issues:
You can also use Freeplay's search functionality to filter the data across any segment of interest:
## Prompt Management (optional)
Freeplay offers [native prompt management](/core-concepts/prompt-management/managing-prompts) , which simplifies the process of version and testing different prompt versions. It allows you to experiment with changes to ADK agent instructions in the Freeplay UI, test different models, and push updates straight to your code, similar to a feature flag.
To leverage Freeplay's prompt management capabilities alongside ADK, you'll want to use the Freeplay ADK agent wrapper. `FreeplayLLMAgent` extends ADK's base `LlmAgent` class, so instead of having to hard code your prompts as agent instructions, you can version prompts in the Freeplay application.
First define a prompt in Freeplay by going to Prompts -> Create prompt template:
When creating your prompt template you'll need to add 3 elements, as described in the following sections:
### System Message
This corresponds to the "instructions" section in your code.
### Agent Context Variable
Adding the following to the bottom of your system message will create a variable for the ongoing agent context to be passed through:
```python python theme={null}
{{agent_context}}
```
### History Block
Click new message and change the role to 'history'. This will ensure the past messages are passed through when present.
Now in your code you can use the `FreeplayLLMAgent`:
```python python theme={null}
from freeplay_python_adk.client import FreeplayADK
from freeplay_python_adk.freeplay_llm_agent import (
FreeplayLLMAgent,
)
FreeplayADK.initialize_observability()
root_agent = FreeplayLLMAgent(
name="social_product_researcher",
tools=[tavily_search],
)
```
When the `social_product_researcher` is invoked, the prompt will be retrieved from Freeplay and formatted with the proper input variables.
## Evaluation
Freeplay enables you to define, version, and run [evaluations](/core-concepts/evaluations/evaluations) from the Freeplay web application. You can define evaluations for any of your prompts or agents by going to Evaluations -> "New evaluation".
These evaluations can be configured to run for both online monitoring and offline evaluation. Datasets for offline evaluation can be uploaded to Freeplay or saved from log examples.
## Dataset Management
As you get data flowing into Freeplay, you can use these logs to start building up [datasets](/core-concepts/datasets/datasets) to test against on a repeated basis. Use production logs to create golden datasets or collections of failure cases that you can use to test against as you make changes.
## Batch Testing
As you iterate on your agent, you can run batch tests (i.e., offline experiments) at both the [prompt](/core-concepts/test-runs/component-level-test-runs) and [end-to-end](/core-concepts/test-runs/end-to-end-test-runs) agent level. This allows you to compare multiple different models or prompt changes and quantify changes head to head across your full agent execution.
[Here](https://github.com/228Labs/freeplay-google-demo/blob/main/examples/example_test_run.py) is a code example for executing a batch test on Freeplay with ADK. [Here](https://github.com/228Labs/freeplay-google-demo/blob/main/examples/example_test_run.py) is a code example for executing a batch test on Freeplay with the Google ADK.
## Sign up now
Go to [Freeplay](https://freeplay.ai/) to sign up for an account, and check out a full Freeplay ADK Integration [here](https://github.com/228Labs/freeplay-google-demo).
# LangGraph
Source: https://docs.freeplay.ai/developer-resources/integrations/langgraph
Add observability, prompt management, and evaluation capabilities to your LangGraph applications.
# Overview
Integrate Freeplay with LangGraph to add observability, prompt management, and evaluation capabilities to your LangGraph applications. This comprehensive guide covers everything from basic setup to advanced agent workflows with state management, streaming, and human-in-the-loop patterns.
## Prerequisites
Before you begin, make sure you have:
* A Freeplay account with an active project
* Python 3.10 or higher installed
* Basic familiarity with LangGraph and LangChain
## Quick Start with Observability
### Installation
Install the Freeplay LangGraph SDK along with your preferred LLM provider. For advanced use, please refer to the documentation on [PyPi](https://pypi.org/project/freeplay-langgraph/).
```bash theme={null}
# Install Freeplay SDK
pip install freeplay-langgraph
# Install your LLM provider (choose one or more)
pip install langchain-openai
pip install langchain-anthropic
pip install langchain-google-vertexai
```
### Configuration
#### Set Up Your Credentials
Configure your Freeplay credentials using environment variables:
```bash theme={null}
export FREEPLAY_API_URL="https://app.freeplay.ai/api"
export FREEPLAY_API_KEY="fp-..."
export FREEPLAY_PROJECT_ID="..."
```
You can find your API key and Project ID in your Freeplay project settings.
#### Initialize the SDK
Create a `FreeplayLangGraph` instance in your application:
```python python theme={null}
from freeplay_langgraph import FreeplayLangGraph
# Using environment variables
freeplay = FreeplayLangGraph()
# Or pass credentials directly
freeplay = FreeplayLangGraph(
freeplay_api_url="https://app.freeplay.ai/api",
freeplay_api_key="fp_...",
project_id="proj_...",
)
```
With this setup, your LangGraph application is now automatically instrumented with OpenTelemetry, sending traces and spans to Freeplay for observability.
Note: It is recommended to manage your prompts within Freeplay to support
better prompt development lifecycle. Continue following this guide to get your
prompts configured within LangGraph.
## Prompt Management
Freeplay's integration requires that you have your prompts configured in Freeplay. By default, `FreeplayLangGraph` fetches prompts from the Freeplay API. This requires you to have prompts configured in Freeplay for use. To learn more, see our Prompt Management guide [here](https://docs.freeplay.ai/core-concepts/prompt-management/managing-prompts). Once configured you will need the prompt names for use in the code.
Managing prompts in Freeplay separates your prompt engineering workflow from your LangGraph application. Instead of hardcoding prompts in your agent code, your team can iterate on prompt templates, test different versions , new models and deploy changes through Freeplay without modifying or redeploying your LangGraph application. This enables your team to test agent behavior, maintain different prompt versions across environments (development, staging, production), and experiment with variations.
### Optional - Prompt Bundling
Once your prompts are saved in Freeplay, you can use bundled prompts stored locally with your application, you can provide a custom template resolver:
```python python theme={null}
from pathlib import Path
from freeplay.resources.prompts import FilesystemTemplateResolver
from freeplay_langgraph import FreeplayLangGraph
# Use filesystem-based prompts bundled with your app
freeplay = FreeplayLangGraph(
template_resolver=FilesystemTemplateResolver(Path("bundled_prompts"))
)
```
This is useful for offline environments, testing, or when you want to version control your prompts alongside your code. See our [Prompt Bundling Guide](https://docs.freeplay.ai/core-concepts/prompt-management/prompt-bundling) to learn more.
## Core Concepts
Freeplay provides two primary ways to work with LangGraph:
1. **`create_agent()`** - For building full LangGraph agents with tool calling, ReAct loops, and state management
2. **`invoke()`** - For simple, stateless LLM invocations when you don't need agent capabilities
Both methods support the same core features: conversation history, tool calling, structured outputs and running tests. Choose `create_agent()` when you need the full power of LangGraph's agent framework, and `invoke()` for simpler use cases.
## Building LangGraph Agents
The `create_agent` method provides full support for LangGraph's agent capabilities including the ReAct loop, tool calling, state management, middleware, and streaming.
### Basic Agent Creation
Create an agent that uses a Freeplay-hosted prompt with automatic model instantiation. You have the ability to pass variables at the creation and invocation of the agent, both are optional depending on your flow:
```python python theme={null}
from freeplay_langgraph import FreeplayLangGraph
from langchain_core.messages import HumanMessage
freeplay = FreeplayLangGraph()
# Create a basic agent with a prmopt stored in Freeplay
agent = freeplay.create_agent(
prompt_name="weather-assistant",
variables={"location": "San Francisco"}, # Optional, enables datasets & testing
environment="production"
)
# Invoke the agent
result = agent.invoke({
"messages": [HumanMessage(content="What's the weather like today?")],
"variables": {{"location": "Denver"}
})
print(result["messages"][-1].content)
```
Using `create_agent` gives you access to LangGraph's full agent capabilities, including tool calling with the ReAct loop, state persistence, and advanced execution control.
### Adding Tools
Bind LangChain tools to your agent for agentic workflows. The agent automatically decides when to call tools:
```python python theme={null}
from langchain_core.tools import tool
@tool
def get_weather(city: str) -> str:
"""Get the current weather for a city."""
return f"Weather in {city}: Sunny, 72°F"
@tool
def get_forecast(city: str, days: int) -> str:
"""Get the weather forecast for a city."""
return f"{days}-day forecast for {city}: Mostly sunny"
agent = freeplay.create_agent(
prompt_name="weather-assistant",
variables={"location": "San Francisco"},
tools=[get_weather, get_forecast],
environment="production"
)
result = agent.invoke({
"messages": [HumanMessage(content="What's the weather in SF and the 5-day forecast?")]
})
```
The agent handles the tool-calling cycle through LangGraph's ReAct loop, deciding when to use tools and when to respond directly to the user.
### Conversation History
Maintain conversation context across multiple turns with conversation history:
```python python theme={null}
from langchain_core.messages import HumanMessage, AIMessage
# Build conversation history
history = [
HumanMessage(content="What's the weather in Paris?"),
AIMessage(content="It's sunny and 22°C in Paris."),
HumanMessage(content="What about in winter?")
]
agent = freeplay.create_agent(
prompt_name="weather-assistant",
variables={"city": "Paris"},
tools=[get_weather],
environment="production"
)
# Pass history in the messages
result = agent.invoke({
"messages": history + [HumanMessage(content="And the average rainfall?")]
})
```
For persistent conversations across multiple invocations, use state persistence with checkpointers (covered in State Management section).
### Structured Output
Get structured, typed responses from your agents using `ToolStrategy` or `ProviderStrategy`:
```python python theme={null}
from pydantic import BaseModel
from langchain.agents.structured_output import ToolStrategy
class WeatherReport(BaseModel):
city: str
temperature: float
conditions: str
humidity: int
agent = freeplay.create_agent(
prompt_name="weather-assistant",
variables={"format": "detailed"},
tools=[get_weather_data],
response_format=ToolStrategy(WeatherReport)
)
result = agent.invoke({
"messages": [HumanMessage(content="Get weather for New York City")]
})
# Access strongly-typed structured output
weather_report = result["structured_response"]
print(f"{weather_report.city}: {weather_report.temperature}°F")
print(f"Conditions: {weather_report.conditions}, Humidity: {weather_report.humidity}%")
```
Structured output ensures your agent returns data in a predictable format, making it easier to integrate with downstream systems, databases, or UIs.
```python python theme={null}
from typing import cast
from langgraph.graph.state import CompiledStateGraph
agent = freeplay.create_agent(...)
# Option 1: Direct unwrap (works at runtime)
state = agent.unwrap().get_state(config)
# Option 2: Cast for full type hints
compiled = cast(CompiledStateGraph, agent.unwrap())
state = compiled.get_state(config) # ✅ Full IDE autocomplete
```
## Automatic Observability
Once initialized, the Freeplay SDK automatically instruments your LangGraph application with OpenTelemetry. This means every LangChain and LangGraph operation is traced and sent to Freeplay without any additional code.
### What Gets Tracked
Freeplay automatically captures:
* **Prompt invocations**: Template, variables, and generated content
* **Model calls**: Provider, model name, tokens used, latency
* **Tool executions**: Which tools were called and their results
* **Agent flows**: Multi-step reasoning and decision paths
* **Conversation flows**: Multi-turn interactions and state transitions
* **Errors and exceptions**: Failed invocations with stack traces
* **Metadata**: Test run IDs, test case IDs, environment names, and custom tags
All metadata is injected automatically through LangChain's `RunnableBindingBase` pattern, ensuring comprehensive observability without manual instrumentation.
### Viewing Traces
You can view all of this data in the Freeplay dashboard, making it easy to:
* Debug issues and understand failure patterns
* Optimize performance and reduce latency
* Understand how your application behaves in production
* Track token usage and costs across environments
* Measure impact of prompt changes over time
## Simple Prompt Invocations
For simpler use cases that don't require the full agent loop, use the `invoke` method. This is ideal for one-off completions, quick classifications, or any scenario where you don't need agent state management or the ReAct loop.
### Basic Invocation
Call a Freeplay-hosted prompt with automatic model instantiation:
```python python theme={null}
from freeplay_langgraph import FreeplayLangGraph
freeplay = FreeplayLangGraph()
# Invoke a prompt - model is automatically created based on Freeplay's config
response = freeplay.invoke(
prompt_name="sentiment-analyzer",
variables={"text": "This product exceeded my expectations!"},
environment="production"
)
print(response.content)
```
Using `invoke` gives you quick access to Freeplay-managed prompts without the overhead of agent state or tool calling. This is perfect for classification tasks, content generation, or any stateless LLM operation.
### Adding Tools
Bind LangChain tools for basic tool calling without the full agent loop:
```python python theme={null}
from langchain_core.tools import tool
@tool
def calculate_discount(price: float, discount_percent: float) -> float:
"""Calculate the final price after applying a discount."""
return price \* (1 - discount_percent / 100)
@tool
def check_inventory(product_id: str) -> int:
"""Check inventory levels for a product."""
return 42 # Mock inventory count
response = freeplay.invoke(
prompt_name="pricing-assistant",
variables={"product": "laptop", "base_price": 1200},
tools=[calculate_discount, check_inventory]
)
```
### Conversation History
Maintain conversation context across multiple turns:
```python python theme={null}
from langchain_core.messages import HumanMessage, AIMessage
# Build conversation history
history = [
HumanMessage(content="What's the weather in Paris?"),
AIMessage(content="It's sunny and 22°C in Paris."),
HumanMessage(content="What about in winter?")
]
# The prompt has full context of the conversation
response = freeplay.invoke(
prompt_name="weather-assistant",
variables={"city": "Paris"},
history=history
)
print(response.content)
```
By passing conversation history, your prompts can maintain context across multiple turns without needing full agent state management.
### Test Execution Tracking
Track test runs for evaluation workflows by pulling test cases from Freeplay and executing them with automatic tracking. By associating invocations with test runs and test cases, you can analyze performance across your test suite, identify regressions, and measure the impact of prompt changes in Freeplay's evaluation dashboard. See more about running end to end test runs [here](https://docs.freeplay.ai/core-concepts/test-runs/end-to-end-test-runs).
#### Creating Test Runs
```python python theme={null}
import os
from freeplay_langgraph import FreeplayLangGraph
from langchain_core.messages import HumanMessage
freeplay = FreeplayLangGraph()
# Create a test run from a dataset
test_run = freeplay.client.test_runs.create(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
testlist="name of the dataset",
name="name your test run",
)
print(f"Created test run: {test_run.id}")
```
#### Executing Test Cases with Simple Invocations
For simple prompt invocations, use the test tracking parameters directly:
```python python theme={null}
# Execute each test case
for test_case in test_run.test_cases:
response = freeplay.invoke(
prompt_name="my-prompt",
variables=test_case.variables,
test_run_id=test_run.id,
test_case_id=test_case.id
)
print(f"Test case {test_case.id}: {response.content}")
```
#### Executing Test Cases with Agents
For LangGraph agents, pass test tracking metadata via config to reuse the agent efficiently:
```python python theme={null}
from langchain_core.messages import HumanMessage
# Create agent once (no test tracking at creation)
agent = freeplay.create_agent(
prompt_name="my-prompt",
variables={"input": "prompt input"},
tools=[get_weather],
)
# Execute each test case with metadata override
for test_case in test_run.trace_test_cases:
result = agent.invoke(
{"messages": [HumanMessage(content=test_case.input)]},
config={
"metadata": {
"freeplay.test_run_id": test_run.id,
"freeplay.test_case_id": test_case.id
}
}
)
print(f"Test case {test_case.id}: {result['messages'][-1].content}")
```
### Using Custom Models
Provide your own pre-configured LangChain model for more control:
```python python theme={null}
from langchain_openai import ChatOpenAI
# Configure your own model with custom parameters
model = ChatOpenAI(
model="gpt-4",
temperature=0.7,
max_tokens=1000
)
response = freeplay.invoke(
prompt_name="content-generator",
variables={"topic": "sustainable energy"},
model=model
)
```
### Async Support
All methods in the Freeplay SDK support async/await for better performance in async applications:
#### Async Agent Invocation
```python python theme={null}
# Async agent creation and invocation
agent = freeplay.create_agent(
prompt_name="assistant",
variables={"role": "helpful"},
tools=[search_knowledge_base]
)
result = await agent.ainvoke({
"messages": [HumanMessage(content="Help me find information")]
})
```
#### Async Simple Invocations
```python python theme={null}
# Async invocation
response = await freeplay.ainvoke(
prompt_name="sentiment-analyzer",
variables={"text": "Great product!"}
)
# Async streaming
async for chunk in freeplay.astream(
prompt_name="content-generator",
variables={"topic": "machine learning"}
):
print(chunk.content, end="", flush=True)
```
#### Async State Management
```python python theme={null}
# Async state inspection
state = await agent.unwrap().aget_state(config)
# Async state updates
await agent.unwrap().aupdate_state(config, {"approval": "granted"})
```
Using async methods improves throughput and reduces latency in applications that handle multiple concurrent requests, such as web servers or API endpoints.
## Supported LLM Providers
Freeplay's LangGraph SDK supports automatic model instantiation for multiple providers. Install the corresponding LangChain integration package for your provider:
### OpenAI
```bash theme={null}
pip install langchain-openai
```
### Anthropic
```bash theme={null}
pip install langchain-anthropic
```
### Vertex AI (Google)
```bash theme={null}
pip install langchain-google-vertexai
```
The SDK automatically detects which provider your Freeplay prompt is configured to use and instantiates the appropriate model with the correct parameters.
# OpenTelemetry
Source: https://docs.freeplay.ai/developer-resources/integrations/tracing-with-otel
Capture LLM observability data using OpenTelemetry for framework-agnostic tracing.
# Overview
Building LLM applications with modern frameworks is great—until you need to understand what's actually happening under the hood. Traditional observability tools weren't built for the nuances of LLM interactions: token counts, model parameters, prompt templates, tool calls, and multi-step agent workflows.
That's where OpenTelemetry (OTel) comes in. By integrating Freeplay with OTel, you get purpose-built LLM observability that works with any framework or orchestration approach. Whether you're using Langgraph, building custom agents, or mixing frameworks, Freeplay captures what matters—automatically.
Freeplay uses OpenTelemetry as a protocol to record LLM observability data. It's not intended to provide general application telemetry for non-LLM related code, so sending arbitrary telemetry to Freeplay will not work. Freeplay supports traces that conform to the OpenInference semantic conventions.
## Why OTel + Freeplay?
* **Framework flexibility**: Your team uses Langgraph. Another team built custom agents. A third is evaluating Google ADK. With OTel, one integration supports all of them.
* **LLM-native insights**: We automatically capture model parameters, token counts, tool schemas, and prompt templates—the data you actually need to improve your AI applications.
* **No architectural changes**: Freeplay observes your orchestration logic without becoming part of it. Your code stays clean, your flexibility stays intact.
* **Built on standards**: OpenInference semantic conventions mean your instrumentation is portable and future-proof.
**Freeplay focuses on LLM observability**
We only record traces and spans containing meaningful LLM information per OpenInference semantic conventions—not general application telemetry.
**Recommended approach:**
1. Use OpenInference instrumentation libraries when available (easiest option)
2. Follow OpenInference semantic conventions for custom instrumentation
3. For advanced control, use our OTel-compliant API directly
If your framework lacks an OpenInference library, you can still record data using standard OpenInference attributes like `input.value`, `output.value`, and `gen_ai.request.model`. See the [supported attributes reference](#supported-attributes-reference) below.
## Getting Started
### Step 1: Install dependencies
```bash theme={null}
pip install freeplay opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp
```
### Step 2: Configure the OTel exporter
Set up Freeplay as your OTel span processing endpoint:
```python python theme={null}
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk import trace as trace_sdk
import os
freeplay_api_url = "https://app.freeplay.ai/api"
exporter = OTLPSpanExporter(
endpoint=f"{freeplay_api_url}/v0/otel/v1/traces",
headers={
"Authorization": f"Bearer {os.environ['FREEPLAY_API_KEY']}",
"X-Freeplay-Project-Id": os.environ["FREEPLAY_PROJECT_ID"],
},
)
tracer_provider = trace_sdk.TracerProvider()
tracer_provider.add_span_processor(SimpleSpanProcessor(exporter))
```
**Getting your credentials:**
* API Key: [Freeplay Settings → API Access](https://app.freeplay.ai/settings/api-access)
* Project ID: Found in your project settings
### Step 3: Instrument your application
We recommend using OpenInference instrumentation libraries when available. For example, with Google ADK:
```python python theme={null}
from google.adk.instrumentation import GoogleADKInstrumentor
GoogleADKInstrumentor().instrument(tracer_provider=tracer_provider)
```
That's it! Your LLM interactions are now being captured and sent to Freeplay.
### Step 4: Enable debugging (optional but recommended)
During development, print spans to your console for easier debugging:
```python python theme={null}
from opentelemetry.sdk.trace.export import ConsoleSpanExporter
tracer_provider.add_span_processor(
SimpleSpanProcessor(ConsoleSpanExporter())
)
```
## Framework examples
### Langgraph/Langchain
### [Langgraph/Langchain Integration](https://colab.research.google.com/drive/1PzB4DoN_YVcwj2RrWas93cH8UOCHdvxG?usp=sharing)
[Complete end-to-end example integrating OTel with Langgraph](https://colab.research.google.com/drive/1PzB4DoN_YVcwj2RrWas93cH8UOCHdvxG?usp=sharing)
### Custom instrumentation
If you're building with a framework that doesn't have an OpenInference instrumentation library, you can manually instrument your code following OpenInference conventions:
```python python theme={null}
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
with tracer.start_as_current_span("llm_call") as span:
# Set OpenInference attributes
span.set_attribute("openinference.span.kind", "LLM")
span.set_attribute("gen_ai.request.model", "gpt-4")
span.set_attribute("input.value", user_query)
# Make your LLM call
response = your_llm_call(user_query)
# Record output and token usage
span.set_attribute("output.value", response)
span.set_attribute("gen_ai.usage.input_tokens", token_count.input)
span.set_attribute("gen_ai.usage.output_tokens", token_count.output)
```
Your span must have a kind value of either `SPAN_KIND_INTERNAL` or `SPAN_KIND_SERVER`.
Your spans must have an attribute named `openinference.span.kind` with a value of `LLM`, `AGENT`, `CHAIN`, or `TOOL`.
Other spans will be ignored.
Freeplay only records data data described below, not general application telemetry.
## Key concepts
### Span types
Freeplay processes different span types based on `openinference.span.kind`:
* **LLM**: Direct LLM API calls with prompts, completions, and token usage
* **AGENT**: Higher-level agentic workflows with decision-making
* **CHAIN**: Sequential operations or pipelines
* **TOOL**: External tool or function calls
Set the appropriate span kind to ensure Freeplay correctly interprets your traces.
### Sessions and traces
Use session and trace identifiers to organize related interactions:
```python python theme={null}
# Link multiple traces to the same conversation session
span.set_attribute("session.id", session_id)
# Or use Freeplay-specific attributes (takes precedence)
span.set_attribute("freeplay.session.id", session_id)
# Name your agent for better organization
span.set_attribute("agent.name", "customer_support_agent")
```
This enables you to track multi-turn conversations and group related agent runs in Freeplay's observability UI.
### Environment attribute
You should add environment attribute to your spans that corresponds with the
prompt template environment that you are using. For example, if you are using a
prompt template in the "production" environment, you should set the environment
attribute to "production".
```python python theme={null}
span.set_attribute("freeplay.environment", "production")
# or "staging", "development", etc.
```
## Supported attributes reference
### OpenInference attributes
Freeplay maps OpenInference attributes to internal fields for consistent observability. These are the standard attributes you should use when instrumenting your LLM applications:
| OTEL Attribute | Freeplay Field | Payload Type | Notes |
| ------------------------------- | -------------------------------------------- | ------------- | -------------------------------------------------- |
| `openinference.span.kind` | N/A | RecordPayload | Determines span type (LLM, AGENT, CHAIN, TOOL) |
| `gen_ai.request.model` | `completion.model_name` | RecordPayload | LLM model name |
| `llm.model_name` | `completion.model_name` | RecordPayload | Alternative LLM model name |
| `gen_ai.usage.input_tokens` | `completion.token_counts.prompt_token_count` | RecordPayload | Input token count |
| `llm.token_count.prompt` | `completion.token_counts.prompt_token_count` | RecordPayload | Alternative input token count |
| `gen_ai.usage.output_tokens` | `completion.token_counts.return_token_count` | RecordPayload | Output token count |
| `llm.token_count.completion` | `completion.token_counts.return_token_count` | RecordPayload | Alternative output token count |
| `llm.provider` | `completion.provider_name` | RecordPayload | LLM provider (openai, anthropic, google) |
| `gen_ai.system` | `completion.provider_name` | RecordPayload | Alternative provider identifier |
| `llm.input_messages.*` | `completion.api_messages` | RecordPayload | Input messages (flattened array format) |
| `llm.output_messages.*` | `completion.api_messages` | RecordPayload | Output messages (flattened array format) |
| `llm.tools.*` | `completion.tool_schema` | RecordPayload | Tool/function definitions |
| `llm.prompt_template.variables` | `completion.inputs` | RecordPayload | Variables applied to template (merged into inputs) |
| `llm.invocation_parameters` | `completion.llm_parameters` | RecordPayload | Parameters used during LLM invocation (JSON) |
| `input.value` | `trace.input` | TracePayload | Trace input data |
| `output.value` | `trace.output` | TracePayload | Trace output data |
| `tool.name` | `trace.name` | TracePayload | Tool span name |
| `agent.name` | `trace.agent_name` | TracePayload | Agent span name |
| `session.id` | `completion.session_id` | RecordPayload | Session identifier |
| `session.id` | `trace.session_id` | TracePayload | Session identifier |
### Freeplay-specific attributes
Enhance your traces with Freeplay-specific metadata for tighter integration with Freeplay features:
| OTEL Attribute | Freeplay Field | Payload Type | Notes |
| ------------------------------------- | --------------------------------------- | ------------- | -------------------------------------------------------- |
| `freeplay.session.id` | `completion.session_id` | RecordPayload | Freeplay session ID (takes precedence over `session.id`) |
| `freeplay.session.id` | `trace.session_id` | TracePayload | Freeplay session ID (takes precedence over `session.id`) |
| `freeplay.environment` | `completion.tag` | RecordPayload | Environment tag (e.g., "production", "staging") |
| `freeplay.prompt_template.id` | `completion.prompt_template_id` | RecordPayload | Link to Freeplay prompt template |
| `freeplay.prompt_template.version.id` | `completion.prompt_template_version_id` | RecordPayload | Specific prompt template version |
| `freeplay.test_run.id` | `completion.test_run_id` | RecordPayload | Associate with test run |
| `freeplay.test_case.id` | `completion.test_case_id` | RecordPayload | Associate with specific test case |
| `freeplay.input_variables` | `completion.inputs` | RecordPayload | Input variables (JSON) |
| `freeplay.metadata.*` | `completion.custom_metadata` | RecordPayload | Custom metadata (flattened dict format) |
| `freeplay.eval_results.*` | `completion.eval_results` | RecordPayload | Evaluation results (flattened dict format) |
## Next steps
Once you've integrated OTel with Freeplay, you can:
* **View traces** in the [Freeplay observability dashboard](https://app.freeplay.ai)
* **Run evaluations** on your logged interactions to measure quality
* **Build datasets** from production traces for systematic testing
* **Monitor performance** across different model versions and configurations
* **Track costs** with automatic token usage and cost calculations
## Additional resources
* [OpenInference Semantic Conventions](https://github.com/Arize-ai/openinference)
* [OpenTelemetry Documentation](https://opentelemetry.io/docs/)
* [Freeplay Sessions, Traces, and Completions Guide](/core-concepts/observability/sessions-traces-and-completions)
***
# Vercel AI SDK
Source: https://docs.freeplay.ai/developer-resources/integrations/vercel-ai-sdk
Build a powerful and well evaluated agent with Freeplay and the Vercel AI SDK
# Overview
Vercel's AI SDK is a robust, end to end framework for powering your Agent. Freeplay provides a simple, lightweight integration to instrument your Agent, making it easy for you to:
* Log agent interactions
* Evaluate your agent's behavior and identify issues
* Iterate on your system prompt and monitor the impact
## This guide will walk you through:
1. **Initializing Freeplay in your code** - Installing dependencies and configuring your environment
2. **Managing your prompts in Freeplay** - Migrating prompt and model configurations to Freeplay and accessing them in your app
3. **Integrating observability** - Adding Freeplay logging to your SDK code for comprehensive data capture and agent tracing
## Prerequisites
Before you begin, make sure you have
* A Freeplay account with an active project
* You're using Vercel AI SDK v5 or greater
## Installation
Install the Freeplay Vercel SDK along with the required dependencies
```bash theme={null}
npm install @freeplayai/vercel @vercel/otel @arizeai/openinference-vercel @opentelemetry/api @opentelemetry/sdk-trace-base
```
(Note: If you're using the NextJS implementation, you can omit `@opentelemetry/sdk-trace-base`)
## Configuration
### Set up your credentials
Configure the following environment variables:
```bash theme={null}
# Freeplay credentials
FREEPLAY_API_KEY=your_freeplay_api_key_here
FREEPLAY_PROJECT_ID=your_freeplay_project_id_here
# IMPORTANT: If you're on a custom deployment, like "acme.freeplay.ai", specify an OTEL endpoint.
# Default is "https://app.freeplay.ai/api/v0/otel/v1/traces"
FREEPLAY_OTEL_ENDPOINT=https://acme.freeplay.ai/api/v0/otel/v1/traces
# Allowed provider API keys, select at least one:
# AI Gateway - Model still must be OpenAI, Anthropic, Google or Vertex
AI_GATEWAY_API_KEY=
# OpenAI
OPENAI_API_KEY=
# Anthropic
ANTHROPIC_API_KEY=
# Google
GOOGLE_GENERATIVE_AI_API_KEY=
# Google Vertex
GOOGLE_VERTEX_LOCATION=
GOOGLE_VERTEX_PROJECT=
GOOGLE_CLIENT_EMAIL=
GOOGLE_PRIVATE_KEY=
GOOGLE_PRIVATE_KEY_ID=
```
You can find your API key and Project ID in your Freeplay project settings.
### Initialize the `FreeplaySpanProcessor`
#### Next JS `instrumentation.ts`
If you're using Next's included Telemetry harness, you can simply add the Freeplay processor to your existing setup
```typescript typescript theme={null}
import { registerOTel } from "@vercel/otel";
import { createFreeplaySpanProcessor } from "@freeplayai/vercel";
export function register() {
registerOTel({
serviceName: "otel-nextjs-example",
spanProcessors: [createFreeplaySpanProcessor(), ...otherProcessors],
});
}
```
#### Node manual telemetry setup
If you're not using Vercel's Telemetry, you can add Freeplay's Span Processor to another library when you initialize your app. For example:
```typescript typescript theme={null}
import { NodeSDK } from "@opentelemetry/sdk-node";
// Initialize OpenTelemetry with Freeplay
const sdk = new NodeSDK({
spanProcessors: [createFreeplaySpanProcessor()],
});
```
### Using Freeplay-Hosted Prompts (Recommended)
One of Freeplay's most powerful features is centralized prompt management. Instead of hardcoding prompts in your application, call them from Freeplay with version control and environment management.
```javascript theme={null}
import { streamText } from "ai";
import {
getPrompt,
FreeplayModel,
createFreeplayTelemetry,
} from "@freeplayai/vercel";
export async function POST(req: Request) {
const { messages, chatId } = await req.json();
const inputVariables = {
customer_issue: "I can't log into my account",
};
// Get prompt from Freeplay
const prompt = await getPrompt({
templateName: "customer-support-agent", // Replace with your prompt name
variables: inputVariables,
messages,
});
// Automatically select the correct model provider based on the prompt
const model = await FreeplayModel(prompt);
const result = streamText({
model,
messages,
system: prompt.systemContent,
experimental_telemetry: createFreeplayTelemetry(prompt, {
functionId: "my-streamText-agent",
sessionId: chatId,
inputVariables,
}),
});
return result.toDataStreamResponse();
}
```
### Using Raw OTEL
With this method, you can use the AI SDK as you normally may, and just need to add the following code snippet to your `streamText` (or similar) method you're using for chat.
NOTE: This is not recommended except for initial testing, as many Freeplay features will not be available if you do not implement Freeplay prompt management.
```javascript theme={null}
experimental_telemetry: {
isEnabled: true,
functionId: "my-streamText-agent",
metadata: {
sessionId: chatId,
},
},
```
## Automatic Observability
Once initialized, the Freeplay SDK automatically instruments your Vercel AI SDK application with OpenTelemetry. This means every chat turn is traced and sent to Freeplay without any additional code.
### What Gets Tracked
Freeplay automatically captures:
* **LLM Interactions**: Prompt version, inputs, provider, model name, tokens used, latency
* **Tool executions**: Which tools were called and their results
* **Conversation flows**: Multi-turn interactions and state transitions
You can view all of this data in the Freeplay dashboard, making it easy to debug issues, optimize performance, and understand how your application behaves in production.
## Next Steps
Now that you've integrated Freeplay with your AI SDK application, you can:
* **Create and manage prompts** in the Freeplay dashboard with version control
* **Set up environments** to test changes in staging before deploying to production
* **Build evaluation datasets** to systematically test your application's performance
* **Analyze traces** to identify bottlenecks and optimize your agent workflows
* **Collaborate with your team** on prompt engineering and application improvements
Visit the [Freeplay documentation](https://docs.freeplay.ai) to learn more about advanced features like prompt experiments, A/B testing, and custom evaluation metrics.
***
# Developer Resources
Source: https://docs.freeplay.ai/developer-resources/overview
SDKs, integrations, and APIs for building with Freeplay.
Freeplay provides multiple integration paths depending on your stack and needs. This page helps you understand how they fit together and choose the right approach.
## Integration Philosophy
Freeplay follows a layered approach:
1. **HTTP API** - The foundation. All Freeplay functionality is accessible via REST endpoints.
2. **Native SDKs** - Language-specific bindings for common operations (Python, TypeScript, Java/Kotlin).
3. **Framework Integrations** - Packages optimized for specific AI frameworks (LangGraph, Vercel AI SDK, Google ADK).
4. **OpenTelemetry** - For observability with OTel-compatible frameworks where Freeplay lacks a direct integration.
The SDKs are designed for core Freeplay functionality that is likely in your production code path: fetching prompts, recording completions, and executing tests. The API provides a superset of functionality for automation, bulk operations, and advanced use cases. OpenTelemetry provides observability only—use it alongside the SDK or API for prompt management.
## Freeplay SDKs
Native SDKs for direct integration with full control over prompt management, observability, and testing / evaluation.
Full-featured SDK for Python applications
Native support for Node.js and TypeScript projects
JVM SDK for Java and Kotlin applications
**What the SDKs provide:**
* Fetch and format prompt templates with variable interpolation
* Record completions, traces, and sessions for observability
* Execute batch tests using saved datasets
* Add customer feedback to observability logs
[View SDK documentation →](/freeplay-sdk/organizing-principles)
## AI Framework Integrations
For teams using popular AI frameworks, dedicated integration packages provide simplified, automatic observability and streamlined prompt management.
| Integration | Language | Observability | Prompt Management | Best For |
| -------------------------------------------------------------------- | ----------- | ----------------- | ----------------------- | ----------------------------------- |
| [LangGraph](/developer-resources/integrations/langgraph) | Python only | Automatic tracing | Full support | LangGraph agents |
| [Vercel AI SDK](/developer-resources/integrations/vercel-ai-sdk) | TypeScript | Automatic tracing | Full support | TypeScript/JS AI applications |
| [Google ADK](/developer-resources/integrations/adk) | Python | Automatic tracing | Full support | Google ADK agents |
| [OpenTelemetry](/developer-resources/integrations/tracing-with-otel) | Any | "AI centric" OTel | Not included (use SDKs) | Vendor agnostic / custom frameworks |
OpenTelemetry integration provides observability only. For prompt management with OTel-traced applications, use the Freeplay SDK alongside your OTel instrumentation.
## HTTP API
The REST API provides programmatic access to all Freeplay capabilities. Use the API when you need:
* Operations not covered by the SDK (e.g. bulk uploads, search, statistics)
* Integration with languages without a native SDK
* Additional automation and scripting
Complete HTTP API documentation with authentication, endpoints, and interactive playground
## Code Recipes
Complete, runnable examples for common integration patterns. Use these as starting points.
Browse examples for prompts, chat, tool calling, providers, and testing
## Choosing Your Integration
| Your Situation | Recommended Path |
| ------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| New, custom application in Python/TypeScript/Java | Start with the [Freeplay SDK](/freeplay-sdk/setup) |
| New, custom application in other languages (Go, Ruby, etc.) | Use the [HTTP API](/api-reference) directly |
| Building with LangGraph in Python | Use the [LangGraph integration](/developer-resources/integrations/langgraph) |
| Using Vercel AI SDK | Use the [Vercel AI SDK integration](/developer-resources/integrations/vercel-ai-sdk) |
| Using Google ADK | Use the [ADK integration](/developer-resources/integrations/adk) |
| Using another OTel-compatible framework (LlamaIndex, etc.) or prefer OTel | [OpenTelemetry](/developer-resources/integrations/tracing-with-otel) for observability + SDK for prompts |
| Need additional bulk operations or automation | [HTTP API](/api-reference) directly |
| Want working examples | Browse [Code Recipes](/developer-resources/recipes/overview) |
## MCP integration & Freeplay Skills
These tools are experimental and subject to change. Use only with trusted agents, as they provide access to your Freeplay API credentials.
For teams using Claude Code, Cursor, Claude Desktop, or similar, Freeplay provides experimental integrations that enable AI agents to interact directly with your Freeplay workspace through natural language.
Model Context Protocol server with tools and skills
Specialized skills for Claude Code and Cursor
Native Claude Code plugin (bundles MCP server)
What these tools enable:
* Analyze production logs and diagnose quality issues through conversation
* Iterate on prompts and agent configurations using real production data
* Run experiments and manage datasets directly from your editor
* Debug AI systems by exploring traces and sessions interactively
These integrations are ideal for development and debugging workflows, allowing you to explore Freeplay data and iterate on AI systems conversationally. For production integrations, use the SDKs or HTTP API above. See the respective GitHub repositories for installation instructions.
## Production Best Practices
Many Freeplay customers configure different client setups for different environments:
* **Dev/Staging**: Fetch prompts from the Freeplay server for rapid iteration
* **Production**: Use [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling) to read prompts from local files
This provides fast experimentation in lower environments while ensuring production stability. See [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling) for implementation details.
# Call Anthropic on Bedrock
Source: https://docs.freeplay.ai/developer-resources/recipes/call-anthropic-on-bedrock
Call Anthropic models via AWS Bedrock with Freeplay integration.
### 1. Configure Bedrock in Freeplay App
Configure Bedrock as a provider in your Freeplay account by going to
Models -> Toggle on Bedrock ->
Deploy a prompt with Bedrock
### 2. Import Anthropic Bedrock Client
Anthropic offers a specific Bedrock client in their SDK
### 3. Create a Bedrock Client
Instantiate your Bedrock client using the anthropic SDK
### 4. Fetch and Format Prompt
Retrieve a formatted prompt from freeplay
### 5. Call Bedrock Model
Call your Bedrock model
### 6. Record to Freeplay
Record the interaction back to freeplay
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
from anthropic.lib.bedrock import AnthropicBedrock
from anthropic import NotGiven
import os
import time
from dotenv import load_dotenv
load_dotenv()
project_id = os.getenv("FREEPLAY_PROJECT_ID")
freeplay_key = os.getenv("FREEPLAY_API_KEY")
freeplay_url = os.getenv("FREEPLAY_URL")
aws_access_key_id = os.getenv("AWS_ACCESS_KEY_ID")
aws_secret_key = os.getenv("AWS_SECRET_KEY")
fpClient = Freeplay(
api_base=freeplay_url,
freeplay_api_key=freeplay_key
)
bedrockClient = AnthropicBedrock(
aws_secret_key=aws_secret_key,
aws_access_key=aws_access_key_id,
aws_region="us-east-1"
)
# fetch the prompt
prompt_vars = {
"pop_star": "Taylor Swift",
}
formatted_prompt = fpClient.prompts.get_formatted(
project_id=project_id,
template_name="album_bot",
environment="latest",
variables=prompt_vars
)
s = time.time()
response = bedrockClient.messages.create(
model=formatted_prompt.prompt_info.model,
system=formatted_prompt.system_content or NotGiven(),
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
e = time.time()
response_text = response.content[0].text
all_messages = formatted_prompt.all_messages(
{"role": "assistant",
"content": response_text}
)
session = fpClient.sessions.create()
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=s, end_time=e)
)
fpClient.recordings.create(payload)
```
# Call Llama 3 on AWS SageMaker
Source: https://docs.freeplay.ai/developer-resources/recipes/call-llama-3-on-aws-sagemaker
Call Llama models hosted on AWS SageMaker with Freeplay prompt management.
### 1. Configure Sagemaker Endpoint
Configure your Sagemaker endpoint in the Freeplay app and set your prompt to use it
[https://docs.freeplay.ai/account-setup/model-management#model-management](https://docs.freeplay.ai/account-setup/model-management#model-management)
### 2. Configure your Freeplay Client
### 3. Configure your Sagemaker Client
Configure your Sagemaker client using your AWS access keys
### 4. Fetch and Format your Prompt
Fetch your prompt from Freeplay, formatting it with your input variables
### 5. Call your Sagemaker Endpoint
Call your sagemaker endpoint directly using the formatted prompt object to key off needed parameters
### 6. Record to Freeplay
Record the interaction back to freeplay!
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, TestRunInfo, CallInfo
import boto3
import os
from dotenv import load_dotenv
import time
import json
load_dotenv("../.env")
project_id = os.getenv("FREEPLAY_PROJECT_ID")
freeplay_key = os.getenv("FREEPLAY_API_KEY")
freeplay_url = os.getenv("FREEPLAY_URL")
aws_access_key_id = os.getenv("AWS_ACCESS_KEY_ID")
aws_secret_key = os.getenv("AWS_SECRET_KEY")
fpClient = Freeplay(
api_base=freeplay_url,
freeplay_api_key=freeplay_key
)
sagemakerClient = boto3.client(
'sagemaker-runtime',
region_name='us-east-1',
aws_access_key_id=aws_access_key_id,
aws_secret_access_key=aws_secret_key
)
# fetch the prompt
prompt_vars = {
"pop_star": "Taylor Swift",
}
formatted_prompt = fpClient.prompts.get_formatted(
project_id=project_id,
template_name="album_bot",
environment="latest",
variables=prompt_vars
)
endpoint_name = formatted_prompt.prompt_info.provider_info['endpoint_name']
inference_component_name = formatted_prompt.prompt_info.provider_info['inference_component_name']
payload_body = {
"inputs": formatted_prompt.llm_prompt_text,
"parameters": formatted_prompt.prompt_info.model_parameters
}
# make the llm call
start = time.time()
response = sagemakerClient.invoke_endpoint(
EndpointName=endpoint_name,
InferenceComponentName=inference_component_name,
ContentType='application/json',
Accept='application/json',
Body=json.dumps(payload_body)
)
end = time.time()
response_content = json.loads(response['Body'].read().decode('utf-8'))['generated_text']
print("response_content: ", response_content)
all_messages = formatted_prompt.all_messages(
{"role": "assistant",
"content": response_content}
)
session = fpClient.sessions.create()
fpClient.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=start, end_time=end)
)
)
```
# Call OpenAI on Azure
Source: https://docs.freeplay.ai/developer-resources/recipes/call-openai-on-azure
Call OpenAI models via Azure OpenAI Service with Freeplay integration.
### 1. Set Up
* Make sure you have your Azure Endpoint configured in Freeplay. See Using Freeplay -> Model Management for more details
* Load environment variables
* configure Freeplay client
### 2. Fetch your Prompt
Fetch and format your Prompt Template from Freeplay
### 3. Configure Azure Client
All details you need to configure your Azure OpenAI client can be found in the provider\_info section of your FormattedPrompt object
### 4. Make your LLM Call
Make your call directly to Azure OpenAI keying all details off of your Formatted Prompt object
### 5. Record to Freeplay
Update your message set and record the interaction back to Freeplay
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo
import openai
from dotenv import load_dotenv
import os
import time
load_dotenv("../.env")
project_id = os.getenv("FREEPLAY_PROJECT_ID")
freeplay_key = os.getenv("FREEPLAY_API_KEY")
freeplay_api_base = os.getenv("FREEPLAY_URL")
azure_api_key = os.getenv("AZURE_API_KEY")
API_VERSION_STRING = '2024-02-15-preview'
# instantiate freeplay client
fpClient = Freeplay(
freeplay_api_key=freeplay_key,
api_base=freeplay_api_base,
)
# get the formatted prompt
prompt_vars = {"feature": "eyes", "celebrity": "Oprah Whinfrey"}
formatted_prompt = fpClient.prompts.get_formatted(project_id=project_id,
template_name="complement",
environment="latest",
variables=prompt_vars)
# configure the azure openai client
azureClient = openai.AzureOpenAI(
api_key=azure_api_key,
api_version=API_VERSION_STRING,
# key your provider info from the prompt template including: Endpoint, Deploy ID and Model
**formatted_prompt.prompt_info.provider_info
)
# make your llm call
s = time.time()
completion = azureClient.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt, # Note: casting may be required for formatting: cast(List[ChatCompletionMessageParam], formatted_prompt.llm_prompt)
**formatted_prompt.prompt_info.model_parameters
)
e = time.time()
# Record to freeplay
# update your messages
all_messages = formatted_prompt.all_messages(
{'role': completion.choices[0].message.role,
'content': completion.choices[0].message.content}
)
# create a session
session = fpClient.sessions.create()
# record your llm call to freeplay
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=s, end_time=e)
)
completion_info = fpClient.recordings.create(payload)
```
# Managing Multi-Turn Chat History
Source: https://docs.freeplay.ai/developer-resources/recipes/continuous-chat
Manage conversation history for multi-turn chat applications with Freeplay.
### 1. Instantiate Clients
Instantiate Clients for Freeplay and your LLM Provider
### 2. Fetch Prompt Template
Fetch prompt template from Freeplay
### 3. Format Prompt
format prompt including input variables and history
### 4. Call LLM
Call your LLM provider, history will be merged into your prompt and ready to pass through
### 5. Record Interaction
Record the interaction to Freeplay
### 6. Manage history
Determine what to include in conversation history over each turn. Append the select messages to an array
## Examples
```python Python theme={null}
import json
import os
import time
from copy import deepcopy
from typing import Optional
import boto3
from anthropic import Anthropic, NotGiven
from openai import OpenAI
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo, TraceInfo
fp_client = Freeplay(
freeplay_api_key=os.environ['FREEPLAY_API_KEY'],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api"
)
project_id = os.environ['FREEPLAY_PROJECT_ID']
environment = 'dev'
anthropic_client = Anthropic(
api_key=os.environ.get("ANTHROPIC_API_KEY")
)
articles = [
"george washington was the first president of the united states",
"the sky is blue",
"the earth is round",
""
]
questions = [
"who was the first president of the united states?",
"what color is the sky?",
"what shape is the earth?",
"repeat the first question and answer"
]
input_pairs = list(zip(articles, questions))
template_prompt = fp_client.prompts.get(
project_id=project_id,
template_name='History-Basics',
environment=environment
)
def call_and_record(
project_id: str,
template_name: str,
env: str,
history: list,
input_variables: dict,
session_info: SessionInfo,
trace_info: Optional[TraceInfo] = None
) -> dict:
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name=template_name,
environment=env,
variables=input_variables,
history=history,
)
start = time.time()
completion = anthropic_client.messages.create(
system=formatted_prompt.system_content or NotGiven(),
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
llm_response = completion.content[0].text
print("Completion: %s" % llm_response)
assistant_response = {'role': 'assistant', 'content': llm_response}
all_messages = formatted_prompt.all_messages(new_message=assistant_response)
call_info = CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end)
record_response = fp_client.recordings.create(
RecordPayload(
all_messages=all_messages,
session_info=session_info,
inputs=input_variables,
prompt_info=formatted_prompt.prompt_info,
call_info=call_info,
trace_info=trace_info
)
)
return {'completion_id': record_response.completion_id,
'llm_response': assistant_response,
"all_messages": all_messages}
session = fp_client.sessions.create()
history = []
for inputs in input_pairs:
input_vars = {'question': inputs[1], 'article': inputs[0]}
record_response = call_and_record(
project_id=project_id,
template_name='History-QA',
env=environment,
history=history,
input_variables=input_vars,
session_info=session.session_info,
)
history = [msg for msg in record_response['all_messages'] if msg['role'] != 'system']
```
```javascript Node theme={null}
import Freeplay, {getCallInfo, getSessionInfo} from "freeplay";
import Anthropic from "@anthropic-ai/sdk";
const projectId = process.env['FREEPLAY_PROJECT_ID'];
const environment = 'dev';
const templateName = "History-QA";
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: `${process.env["FREEPLAY_API_URL"]}/api`,
});
const anthropicClient = new Anthropic({apiKey: process.env['ANTHROPIC_API_KEY']})
async function call(
projectId,
templateName,
environment,
input_variables,
history,
session_info,
trace_info
){
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName,
environment,
variables: input_variables,
history: history
});
console.log("Prompt", formattedPrompt.llmPrompt);
let start = new Date();
const llmResponse = await anthropicClient.messages.create(
{
model: formattedPrompt.promptInfo.model,
messages: formattedPrompt.llmPrompt,
system: formattedPrompt.systemContent,
...formattedPrompt.promptInfo.modelParameters
}
);
let end = new Date();
const llmResponseText = llmResponse.content[0].text;
let messages = formattedPrompt.allMessages(
{
content: llmResponseText,
role: 'Assistant'
});
const completionResponse = await fpClient.recordings.create(
{
projectId,
allMessages: messages,
inputs: input_variables,
sessionInfo: session_info,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
traceInfo: trace_info
}
)
return {completionId: completionResponse.completionId, userMessages: formattedPrompt.llmPrompt, llmResponseText: llmResponseText};
}
const articles = [
"george washington was the first president of the united states",
"the sky is blue",
"the earth is round",
""
];
const questions = [
"who was the first president of the united states?",
"what color is the sky?",
"what shape is the earth?",
"repeat the first question and answer"
];
const inputPairs = articles.map((article, index) => [article, questions[index]]);
async function main() {
const session = await fpClient.sessions.create();
const history_messages = [];
for (const [article, question] of inputPairs) {
const traceInfo = await session.createTrace(question);
const botResponse = await call(
projectId, templateName, environment,
{ question: question, article: article }, history_messages, getSessionInfo(session), traceInfo
);
// update history
history_messages.push(...botResponse.userMessages);
history_messages.push({
content: botResponse.llmResponseText,
role: 'assistant'
});
console.log("Bot response: ", botResponse.llmResponseText);
await traceInfo.recordOutput(projectId, botResponse.llmResponseText);
}
}
main().catch(console.error);
```
```java Kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
import ai.freeplay.example.java.ThinExampleUtils.callAnthropic // this is a private helper function
import com.fasterxml.jackson.databind.ObjectMapper
import kotlinx.coroutines.future.await
import kotlinx.coroutines.runBlocking
private val objectMapper = ObjectMapper()
fun main(): Unit = runBlocking {
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val baseUrl = System.getenv("FREEPLAY_API_URL") + "/api"
val anthropicApiKey = System.getenv("ANTHROPIC_API_KEY")
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.baseUrl(baseUrl)
)
val questions = listOf(
"who was the first president of the united states?",
"what color is the sky?",
"what shape is the earth?",
"repeat the first question and answer"
)
val articles = listOf(
"george washington was the first president of the united states",
"the sky is blue",
"the earth is round",
""
)
val history = mutableListOf()
println("Getting the prompt...")
val template = fpClient.prompts()
.get(
projectId,
"History-QA",
"latest"
).await()
val sessionInfo = fpClient.sessions().create()
.customMetadata(mapOf("custom_field" to "custom_value"))
.sessionInfo
for (i in 1..questions.size){
val variables = mapOf("question" to questions[i-1], "article" to articles[i-1])
println("variables: $variables")
val formatted = template.bind(variables, history).format>()
println("Calling Anthropic...")
val startTime = System.currentTimeMillis()
val llmResponse = callAnthropic(
objectMapper,
anthropicApiKey,
formatted.promptInfo.model,
formatted.promptInfo.modelParameters,
formatted.formattedPrompt,
formatted.systemContent.orElse(null)
).await()
val bodyNode = objectMapper.readTree(llmResponse.body())
println("Completion: " + bodyNode.path("content").get(0).path("text").asText())
println("Recording the result")
val allMessages: List = formatted.allMessages(
ChatMessage("assistant", bodyNode.path("content").get(0).path("text").asText())
)
if (allMessages.size >= 2) {
history.add(allMessages[allMessages.size - 2])
history.add(allMessages[allMessages.size - 1])
} else if (allMessages.isNotEmpty()) {
history.addAll(allMessages)
}
val callInfo = CallInfo.from(
formatted.promptInfo,
startTime,
System.currentTimeMillis()
)
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace)
).await()
}
}
```
# Pipecat Observer
Source: https://docs.freeplay.ai/developer-resources/recipes/freeplay-pipecat-observer
Add Freeplay observability to Pipecat voice pipelines using the Observer pattern.
### 1. Freeplay & LLM Initialization
Initialize the Freeplay client, get the unformatted prompt (optionally bind variables to it).
Initialize the LLM service for the pipeline using info from Freeplay
### 2. AudioBuffer
Add the audio buffer processor, this allows for us to grab the right audio information to log to Freeplay.
### 3. Add the FreeplayObserver to PipelineTask
### 4. Callbacks to log to Freeplay
Add call backs to handle user and bot audio, when you have the right details, log to Freeplay
### 5. FreeplayObserver
The entire FreeplayObserver class
## Examples
```python Python theme={null}
from helpers.freeplay_observer import FreeplayObserver
from freeplay import Freeplay, SessionInfo
import os
import sys
from dotenv import load_dotenv
from fastapi import WebSocket
from loguru import logger
from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.processors.aggregators.openai_llm_context import OpenAILLMContext
from pipecat.processors.audio.audio_buffer_processor import AudioBufferProcessor
from pipecat.serializers.twilio import TwilioFrameSerializer
from pipecat.services.cartesia.tts import CartesiaTTSService
from pipecat.services.deepgram.stt import DeepgramSTTService
from pipecat.services.openai.llm import OpenAILLMService
from pipecat.transports.network.fastapi_websocket import (
FastAPIWebsocketParams,
FastAPIWebsocketTransport,
)
load_dotenv(override=True)
async def run_bot(
websocket_client: WebSocket,
stream_sid: str,
call_sid: str,
testing: bool,
session: SessionInfo,
):
serializer = TwilioFrameSerializer(
stream_sid=stream_sid,
call_sid=call_sid,
account_sid=os.getenv("TWILIO_ACCOUNT_SID", ""),
auth_token=os.getenv("TWILIO_AUTH_TOKEN", ""),
)
transport = FastAPIWebsocketTransport(
websocket=websocket_client,
params=FastAPIWebsocketParams(
audio_in_enabled=True,
audio_out_enabled=True,
add_wav_header=False,
vad_enabled=True,
vad_analyzer=SileroVADAnalyzer(),
vad_audio_passthrough=True,
serializer=serializer,
),
)
# Example of getting variables from a local function
required_information = select_meeting_details()
# Initialize Freeplay client
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base=os.getenv("FREEPLAY_API_BASE"),
)
# Get the unformatted prompt from Freeplay
unformatted_prompt = fp_client.prompts.get(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=os.getenv("PROMPT_NAME"),
environment="latest",
)
formatted_prompt = unformatted_prompt.bind(
variables={"required_information": required_information},
history=[],
).format()
# Pass the formatted prompt to the LLM
llm = OpenAILLMService(
model=unformatted_prompt.prompt_info.model,
tools=(
unformatted_prompt.tool_schema if unformatted_prompt.tool_schema else None
),
api_key=os.getenv("OPENAI_API_KEY"),
**unformatted_prompt.prompt_info.model_parameters,
)
# Pass the Deepgram API key to the DeepgramSTTService
stt = DeepgramSTTService(
api_key=os.getenv("DEEPGRAM_API_KEY"), audio_passthrough=True
)
tts = CartesiaTTSService(
api_key=os.getenv("CARTESIA_API_KEY"),
voice_id=os.getenv("CARTESIA_VOICE_ID"),
push_silence_after_stop=testing,
)
context = OpenAILLMContext(formatted_prompt.llm_prompt)
context_aggregator = llm.create_context_aggregator(context)
# NOTE: This buffer is what allows us to capture the audio from the user and bot using the callback handlers.
# NOTE: Watch out! This will save all the conversation in memory. You can
# pass `buffer_size` to get periodic callbacks.
audiobuffer = AudioBufferProcessor(
sample_rate=16000, # Optional: desired output sample rate
num_channels=1, # 1 for mono, 2 for stereo
buffer_size=0, # Size in bytes to trigger buffer callbacks
user_continuous_stream=False,
enable_turn_audio=True,
)
freeplay_observer = FreeplayObserver(
fp_client=fp_client,
unformatted_prompt=unformatted_prompt,
environment=os.getenv("FREEPLAY_ENVIRONMENT"),
variables={
"required_information": required_information,
},
)
pipeline = Pipeline(
[
transport.input(), # Websocket input from client
stt, # Speech-To-Text
context_aggregator.user(),
llm, # LLM
tts, # Text-To-Speech
transport.output(), # Websocket output to client
audiobuffer, # Used to buffer the audio in the pipeline
context_aggregator.assistant(),
]
)
task = PipelineTask(
pipeline,
params=PipelineParams(
audio_in_sample_rate=8000,
audio_out_sample_rate=8000,
allow_interruptions=True,
enable_metrics=True,
enable_usage_metrics=True, # This is used to track the usage of the LLM
),
observers=[
freeplay_observer
], # Use the FreeplayObserver to record the audio to Freeplay
)
# Additional handlers to support Freeplay Observer
# save audio bytes from user and store in freeplay_observer
@audiobuffer.event_handler("on_user_turn_audio_data")
async def on_user_turn_audio_data(buffer, audio, sample_rate, num_channels):
if audio and not freeplay_observer._bot_audio:
# aggregate user audio because this event could fire multiple times
# before bot responds
freeplay_observer._user_audio = freeplay_observer._turn_user_audio.extend(
audio
)
freeplay_observer._user_audio = await freeplay_observer.make_wav_bytes(
freeplay_observer._turn_user_audio,
sample_rate,
"user",
prepend_silence_secs=1,
)
elif audio and freeplay_observer._bot_audio:
freeplay_observer._user_audio = freeplay_observer._turn_user_audio.extend(
audio
)
freeplay_observer._user_audio = await freeplay_observer.make_wav_bytes(
freeplay_observer._turn_user_audio,
sample_rate,
"user",
prepend_silence_secs=1,
)
await freeplay_observer.record_to_freeplay()
# save audio bytes from bot and store in freeplay_observer
@audiobuffer.event_handler("on_bot_turn_audio_data")
async def on_bot_turn_audio_data(buffer, audio, sample_rate, num_channels):
# this assumes the user always speaks first and would cut off
# the first turn of the bot
if audio and not freeplay_observer._user_audio:
# aggregate bot audio because this event could fire multiple times
# before user responds
freeplay_observer._bot_audio = await freeplay_observer.make_wav_bytes(
audio, sample_rate, "bot", prepend_silence_secs=1
)
elif audio and freeplay_observer._user_audio:
freeplay_observer._bot_audio = await freeplay_observer.make_wav_bytes(
audio, sample_rate, "bot", prepend_silence_secs=1
)
await freeplay_observer.record_to_freeplay()
##############################################################################
# FreeplayObserver.py
##############################################################################
import os
import io
import wave
import time
import base64
import datetime
import uuid
from pipecat.processors.aggregators.openai_llm_context import OpenAILLMContextFrame
from pipecat.processors.frame_processor import FrameDirection
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo
from dotenv import load_dotenv
import asyncio
import functools
from pipecat.frames.frames import (
LLMFullResponseStartFrame,
LLMFullResponseEndFrame,
TranscriptionFrame,
)
from pipecat.observers.base_observer import BaseObserver, FramePushed
from pipecat.processors.aggregators.llm_response import LLMUserContextAggregator
from pipecat.services.openai.base_llm import BaseOpenAILLMService
load_dotenv(override=True)
class FreeplayObserver(BaseObserver):
def __init__(
self,
fp_client: Freeplay,
unformatted_prompt: str = None,
template_name: str = os.getenv("PROMPT_NAME") or None,
environment: str = "latest",
variables: dict = {},
):
super().__init__()
self.start_llm_interaction = 0
self.end_llm_interaction = 0
self.llm_completion_latency = 0
self.call_id = str(uuid.uuid4()) # Has to be str to record to Freeplay
# Audio related properties
self.sample_width = 2
self.num_channels = 1
self.sample_rate = 16000
self._bot_audio = bytearray()
self._user_audio = bytearray()
self._turn_user_audio = bytearray()
self.user_speaking = False
self.bot_speaking = False
self.fp_client = fp_client
self.session = self.fp_client.sessions.create()
# Freeplay Params
self.template_name = template_name
self.environment = environment
self.unformatted_prompt = unformatted_prompt
self.variables = variables
# Conversation Params
self.conversation_id = self._new_conv_id()
self.conversation_history = []
self.most_recent_user_message = None
self.most_recent_completion = None
def _new_conv_id(self) -> str:
"""Generate a new conversation ID based on the current timestamp (this represents a customer id or similar)."""
return datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
def _reset_recent_messages(self):
"""Reset all temporary message and audio storage."""
self.most_recent_user_message = None
self.most_recent_completion = None
self._user_audio = bytearray()
self._bot_audio = bytearray()
self._turn_user_audio = bytearray()
# self.llm_completion_latency = 0
async def record_to_freeplay(self):
"""Record the current interaction to Freeplay as a new trace."""
# Create a new trace for this interaction
trace = self.session.create_trace(
input=self.most_recent_user_message,
custom_metadata={
"conversation_id": str(self.conversation_id),
},
)
# Add user message to conversation history
self.conversation_history.append(
{
"role": "user",
"content": [
{"type": "text", "text": self.most_recent_user_message},
{
"type": "input_audio",
"input_audio": {
"data": base64.b64encode(self._user_audio).decode("utf-8"),
"format": "wav",
},
},
],
},
)
# Bind the variables to the prompt
if self.unformatted_prompt:
formatted = self.unformatted_prompt.bind(
variables=self.variables,
history=self.conversation_history,
).format()
else:
# Run in executor to avoid blocking
loop = asyncio.get_running_loop()
formatted = await loop.run_in_executor(
None,
self.fp_client.prompts.get_formatted,
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=self.template_name,
environment=self.environment,
history=self.conversation_history,
variables=self.variables,
)
# Calculate latency for the LLM interaction
# Convert nanoseconds to seconds for proper timing
latency_seconds = self.get_llm_response_latency_seconds()
end = time.time()
start = end - latency_seconds
try:
print(f"_____* self._bot_audio: {len(self._bot_audio)}")
# Prepare metadata and record payload
custom_metadata = {"caller_id": self.call_id}
# Prepare assistants response message (mimicing the format of the llm provider message)
assistant_msg = {
"role": "assistant",
"content": [
{"type": "text", "text": self.most_recent_completion},
],
"audio": {
"id": self.conversation_id,
"data": base64.b64encode(self._bot_audio).decode("utf-8"),
"expires_at": 1729234747,
"transcript": self.most_recent_completion,
},
}
# Add assistant's response to conversation history
self.conversation_history.append(assistant_msg)
record = RecordPayload(
project_id=PROJECT_ID,
all_messages=[
*formatted.llm_prompt,
assistant_msg, # Add the assistant's response to the record call
],
session_info=SessionInfo(
self.session.session_id, custom_metadata=custom_metadata
),
inputs={},
prompt_version_info=formatted.prompt_info,
call_info=CallInfo.from_prompt_info(formatted.prompt_info, start, end),
trace_info=trace,
)
# Create recording in Freeplay
loop = asyncio.get_running_loop()
await loop.run_in_executor(None, self.fp_client.recordings.create, record)
# Record output to trace
await loop.run_in_executor(
None,
functools.partial(
trace.record_output,
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
output=self.most_recent_completion,
eval_results={},
),
)
print(
f"✅ Recorded interaction #{len(self.conversation_history) // 2} to Freeplay - LLM response time: {self.get_llm_response_latency_seconds():.3f}s",
flush=True,
)
# Reset only audio and current message data, keep conversation history
self._reset_recent_messages()
except Exception as e:
print(f"❌ Error recording to Freeplay: {e}", flush=True)
# Still reset audio buffers to prevent accumulation
# audio buffers are overwritten in event handler
self._reset_recent_messages()
async def make_wav_bytes(
self, pcm: bytes, sample_rate: int, voice: str, prepend_silence_secs: int = 1
) -> bytes:
"""Convert PCM audio data to WAV format with optional silence prepend."""
if prepend_silence_secs > 0:
silence_samples = int(
self.sample_rate
* self.sample_width
* self.num_channels
* prepend_silence_secs
)
silence = b"\x00" * silence_samples
pcm = silence + pcm
with io.BytesIO() as buf:
with wave.open(buf, "wb") as wf:
wf.setnchannels(self.num_channels)
wf.setsampwidth(self.sample_width)
wf.setframerate(sample_rate)
wf.writeframes(pcm)
return buf.getvalue()
async def on_push_frame(self, data: FramePushed):
src = data.source
dst = data.destination
frame = data.frame
direction = data.direction
timestamp = data.timestamp
# Create direction arrow
arrow = "→" if direction == FrameDirection.DOWNSTREAM else "←"
if isinstance(frame, LLMFullResponseStartFrame):
print(f"LLMFullResponseFrame: START {src} {arrow} {dst}", flush=True)
elif isinstance(frame, LLMFullResponseEndFrame):
print(f"LLMFullResponseFrame: END {src} {arrow} {dst}", flush=True)
elif isinstance(frame, TranscriptionFrame):
# Capture user if bot talks first
if self.most_recent_user_message is None:
self.most_recent_user_message = frame.text
elif isinstance(frame, OpenAILLMContextFrame):
messages = frame.context.messages
# Extract user message and completion from context
# NOTE: this replaces the TranscriptionFrame results, as this maps excatly what the llm recived.
user_messages = [m for m in messages if m.get("role") == "user"]
if user_messages:
self.most_recent_user_message = user_messages[-1].get("content")
completions = [m for m in messages if m.get("role") == "assistant"]
if completions:
self.most_recent_completion = completions[-1].get("content")
# if (
# self.llm_completion_latency
# and self.most_recent_user_message
# and self.most_recent_completion
# ):
# # reset latency
# self.llm_completion_latency = 0
# Get relevant latency metrics for the LLM interaction
if (
isinstance(frame, OpenAILLMContextFrame)
and isinstance(src, LLMUserContextAggregator)
and isinstance(dst, BaseOpenAILLMService)
):
self.start_llm_interaction = timestamp
print(f"_____freeplay-observer.py OpenAILLMContextFrame START: {timestamp}")
elif isinstance(frame, LLMFullResponseEndFrame) and isinstance(
src, BaseOpenAILLMService
):
self.end_llm_interaction = timestamp
# update latency tally
self.llm_completion_latency = (
self.end_llm_interaction - self.start_llm_interaction
)
print(
f"_____freeplay-observer.py * set self.llm_completion_latency: {self.llm_completion_latency} ({self.get_llm_response_latency_seconds():.3f}s)"
)
def get_llm_response_latency_seconds(self):
"""Convert the raw nanosecond LLM response latency to seconds."""
return self.llm_completion_latency / 1_000_000_000
```
# Multi-Chain Prompt with Traces
Source: https://docs.freeplay.ai/developer-resources/recipes/multi-chain-prompt-with-traces
Chain multiple prompts using traces to group related completions.
### 1. Initialization of Freeplay & Clients
### 2. Session & Trace Creation
Create the session and trace, optionally passing additional information to each such as metadata and agent name
### 3. Prompt #1 - Call & Record
### 4. Prompt #2 - Call & Record
### 5. Record the final output to the Trace
Recording the final output to the trace allows you to record evals across the whole trace and the final output (needs to be a str)
## Examples
```python Python theme={null}
import os
import time
from freeplay import Freeplay, CallInfo, RecordPayload
from openai import OpenAI
import re
# Configure environment variables for API access
project_id = os.environ.get("FREEPLAY_PROJECT_ID")
freeplay_api_key = os.environ.get("FREEPLAY_API_KEY")
freeplay_url = os.environ.get("FREEPLAY_URL") # ie "https://app.freeplay.ai/api"
openai_api_key = os.environ.get("OPENAI_API_KEY")
# Initialize OpenAI client
openai = OpenAI(api_key=openai_api_key)
# Initialize Freeplay client with development API endpoint
fp_client = Freeplay(
freeplay_api_key=freeplay_api_key,
api_base=freeplay_url
)
# Create a new Freeplay session to group related completions
session = fp_client.sessions.create({})
user_input = "My favorite artist is Taylor Swift"
# Create a trace to combine multiple prompts into a single workflow
trace_info = session.create_trace(
input=user_input, # Commonly user question but any str input to a trace
agent_name="musicAgent", # Optionally pass an agent name
custom_metadata={
"version": "1.0.8"
}
)
# =============================================================================
# Prompt 1: Generate Album Title
# =============================================================================
# Define variables for the album title generation prompt
prompt_vars_a = {"pop_star": "Taylor Swift"}
# Fetch and format the album generation prompt template
formatted_prompt_a = fp_client.prompts.get_formatted(
project_id=project_id, # Freeplay project identifier
template_name="album_bot", # Name of the prompt template
environment="latest", # Environment tag for prompt versioning
variables=prompt_vars_a # Variables to interpolate into the prompt
)
# Execute the OpenAI completion for album title generation
start = time.time()
chat_completion_a = openai.chat.completions.create(
messages=formatted_prompt_a.llm_prompt,
model=formatted_prompt_a.prompt_info.model,
**formatted_prompt_a.prompt_info.model_parameters
)
end = time.time()
# Extract the generated album name from the response
album_name = chat_completion_a.choices[0].message.content
print(f"Album Name: {album_name}")
# Record the first completion to Freeplay for tracking and analysis
fp_client.recordings.create(
RecordPayload(
all_messages=formatted_prompt_a.all_messages(new_message={
"role": chat_completion_a.choices[0].message.role,
"content": chat_completion_a.choices[0].message.content
}),
inputs=prompt_vars_a,
session_info=session.session_info,
prompt_info=formatted_prompt_a.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt_a.prompt_info, start, end),
trace_info=trace_info
),
)
# =============================================================================
# Prompt 2: Generate Song List for Album
# =============================================================================
# Define variables for the song list generation prompt (includes generated album name)
prompt_vars_b = {"album_name": album_name, "pop_star": "Taylor Swift"}
# Fetch and format the song list generation prompt template
formatted_prompt_b = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="song_bot",
environment="latest",
variables=prompt_vars_b
)
# Execute the OpenAI completion for song list generation
start = time.time()
chat_completion_b = openai.chat.completions.create(
messages=formatted_prompt_b.llm_prompt,
model=formatted_prompt_b.prompt_info.model,
**formatted_prompt_b.prompt_info.model_parameters
)
end = time.time()
song_track = chat_completion_b.choices[0].message.content
# Record the second completion to Freeplay for tracking and analysis
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=formatted_prompt_b.all_messages(new_message={
"role": chat_completion_b.choices[0].message.role,
"content": song_track
}),
inputs=prompt_vars_b,
session_info=session.session_info,
prompt_version_info=formatted_prompt_b.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt_b.prompt_info, start, end),
trace_info=trace_info
)
)
# Record the final output to complete the trace
trace_info.record_output(
project_id=project_id,
output=song_track, # Final trace output (str)
# Optional code evals logged to the trace
eval_results={
"sentiment": 0.7,
"songTrackLength": len(re.findall(r'\n', song_track)) + 1, # count number of lines
}
)
```
# Multi Prompt Chain
Source: https://docs.freeplay.ai/developer-resources/recipes/multi-prompt-chain
Chain multiple prompts together in a session with Freeplay observability.
### 1. Configure your Environment
Configure environment variables and create your freeplay session
### 2. Create a Freeplay Session
Create a Freeplay Session that we will use to tie together all completions in the Chain
### 3. Run the first Completion
* Fetch and format the first prompt
* Call your LLM
* Extract the needed results
### 4. Record the first Completion to Freeplay
* Record to Freeplay
### 5. Run second Completion
* Fetch and format second prompt using output from the first
* Call your LLM provider
* Record to Freeplay
## Examples
```javascript Node theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
import OpenAI from "openai";
import * as dotenv from "dotenv";
// confgure neccesaary environment variables
let projectId = process.env["FREEPLAY_PROJECT_ID"];
let freeplayApiKey = process.env["FREEPLAY_API_KEY"];
let freeplayUrl = process.env["FREEPLAY_URL"];
let openaiApiKey = process.env["OPENAI_API_KEY"];
const openai = new OpenAI(process.env["OPENAI_API_KEY"]);
// instantiate freeplay client
const fpClient = new Freeplay({
freeplayApiKey: freeplayApiKey,
baseUrl: freeplayUrl,
});
// create a freeplay session
// we will tie all completions in the chain to this session
let session = fpClient.sessions.create({});
// fetch the first prompt
// generate an album title for a pop star
const prompt_varsA = {"pop_star": "Taylor Swift"};
const formattedPromptA = await fpClient.prompts.getFormatted({
projectId: projectID,
templateName: "album_bot",
environment: "latest",
variables: {"pop_star": "Taylor Swift"},
});
let start = new Date();
const chatCompletionA = await openai.chat.completions.create({
messages: formattedPromptA.llmPrompt,
model: formattedPromptA.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters
});
let end = new Date();
// get the album name
let album_name = chatCompletionA.choices[0].message.content;
// record to freeplay
fpClient.recordings.create({
allMessages: formattedPromptA.allMessages({
role: chatCompletionA.choices[0].message.role,
content: chatCompletionA.choices[0].message.content,
}),
inputs: prompt_varsA,
sessionInfo: getSessionInfo(session), // tie the completion to the session
promptInfo: formattedPromptA.promptInfo,
callInfo: getCallInfo(formattedPromptA.promptInfo, start, end)
});
// fetch the second prompt
// generate the song list for the album
const prompt_varsB = {"album_name": album_name, pop_star: "Taylor Swift"};
const formattedPromptB = await fpClient.prompts.getFormatted({
projectId: projectID,
templateName: "song_bot",
environment: "latest",
variables: prompt_varsB,
});
start = new Date();
const chatCompletionB = await openai.chat.completions.create({
messages: formattedPromptB.llmPrompt,
model: formattedPromptB.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters
});
end = new Date();
// record to freeplay
fpClient.recordings.create({
projectId,
allMessages: formattedPromptB.allMessages({
role: chatCompletionB.choices[0].message.role,
content: chatCompletionB.choices[0].message.content,
}),
inputs: prompt_varsB,
sessionInfo: getSessionInfo(session), // tie the completion to the session
promptVersionInfo: formattedPromptB.promptInfo,
callInfo: getCallInfo(formattedPromptB.promptInfo, start, end)
});
```
```python Python theme={null}
import os
import time
from freeplay import Freeplay, CallInfo, RecordPayload
from openai import OpenAI
#######################
# Setup #
#######################
# TODO: Use environment variables in production
project_id = os.environ.get("FREEPLAY_PROJECT_ID")
freeplay_api_key = os.environ.get("FREEPLAY_API_KEY")
freeplay_url = os.environ.get("FREEPLAY_URL")
openai_api_key = os.environ.get("OPENAI_API_KEY")
# Initialize clients
openai = OpenAI(api_key=openai_api_key)
fp_client = Freeplay(
freeplay_api_key=freeplay_api_key,
api_base="https://app.freeplay.ai/api"
)
###########################
# Create Freeplay Session #
###########################
# Create a session to group related LLM calls
# This allows tracking the entire chain as a single conversation
session = fp_client.sessions.create()
##################################
# Step 1: Call the first prompt #
##################################
# Define variables for the first prompt
prompt_vars_a = {"pop_star": "Taylor Swift"}
# Get formatted prompt from Freeplay
formatted_prompt_a = fp_client.prompts.get_formatted(
project_id=project_id, # Freeplay project id
template_name="album_bot", # Name of the prompt template
environment="latest", # Version tag of the prompt to use
variables=prompt_vars_a # Variables to populate the template
)
# Execute LLM call with OpenAI
start = time.time()
chat_completion_a = openai.chat.completions.create(
messages=formatted_prompt_a.llm_prompt,
model=formatted_prompt_a.prompt_info.model,
**formatted_prompt_a.prompt_info.model_parameters
)
end = time.time()
# Extract the generated results for use in next step
album_name = chat_completion_a.choices[0].message.content
# Record the completion details back to Freeplay
fp_client.recordings.create(
RecordPayload(
# Include both the original messages and the model's response
all_messages=formatted_prompt_a.all_messages(
new_message={
"role": chat_completion_a.choices[0].message.role,
"content": chat_completion_a.choices[0].message.content
}
),
inputs=prompt_vars_a, # Variables used in the prompt
session_info=session.session_info, # Link to the chain session
prompt_info=formatted_prompt_a.prompt_info, # Prompt metadata
call_info=CallInfo.from_prompt_info( # Timing and performance data
formatted_prompt_a.prompt_info, start, end
)
)
)
#############################################
# Step 2: Call the next prompt in the chain #
#############################################
# Use the album name from step 1 as input for step 2
prompt_vars_b = {
"album_name": album_name,
"pop_star": "Taylor Swift"
}
# Get formatted prompt for song generation
formatted_prompt_b = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="song_bot",
environment="latest",
variables=prompt_vars_b
)
# Execute second LLM call
start = time.time()
chat_completion_b = openai.chat.completions.create(
messages=formatted_prompt_b.llm_prompt,
model=formatted_prompt_b.prompt_info.model,
**formatted_prompt_b.prompt_info.model_parameters
)
end = time.time()
# Record the second completion to Freeplay
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=formatted_prompt_b.all_messages(
new_message={
"role": chat_completion_b.choices[0].message.role,
"content": chat_completion_b.choices[0].message.content
}
),
inputs=prompt_vars_b,
session_info=session.session_info, # Same session as first call
prompt_version_info=formatted_prompt_b.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted_prompt_b.prompt_info, start, end
)
)
)
```
# OpenAI Batch API Example
Source: https://docs.freeplay.ai/developer-resources/recipes/openai-batch-api
Process multiple LLM requests using OpenAI's Batch API with Freeplay.
### 1. Set up
### 2. Batch data
### 3. Fetch Freeplay prompt
### 4. Prepare batch requests
Make sure to include the call to Freeplay, fp\_client.recordings.create(RecordPayload(RecordPayload))
### 5. Pass api\_style to CallInfo
### 6. Perform all batch requests
### 7. Update the Freeplay completions
## Examples
```python Python theme={null}
import json
import os
from openai import OpenAI
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo, TraceInfo
import time
from freeplay.resources.recordings import RecordUpdatePayload
fp_client = Freeplay(
freeplay_api_key=os.environ['FREEPLAY_API_KEY'],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api"
)
project_id = ''
environment = 'latest'
openai_client = OpenAI(
api_key=os.environ.get("OPENAI_API_KEY")
)
questions = [
"What is the capital of France?",
"Who is the president of the United States?",
"What is the population of Tokyo?",
"Who was the star of the movie 'The Matrix'?",
"What is the capital of Japan?",
"What is the highest mountain in the world?",
"Who is the author of 'To Kill a Mockingbird'?",
]
## Create the batch file ##
# fetch a prompt template from freeplay
prompt_template = fp_client.prompts.get(
project_id=project_id,
template_name="basic_trivia_bot",
environment=environment,
)
# loop over each input and create a completion in freeplay as well as a line in the batch file
batch_file_data = []
for question in questions:
# format the prompt with the input
input_vars = {
'question': question,
}
formatted_prompt = prompt_template.bind(input_vars).format()
# create the completion in freeplay in order to get a completion id
session_id = fp_client.sessions.create()
completion_info = fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=formatted_prompt.messages,
inputs=input_vars,
session_info=session_id,
prompt_info=prompt_template.prompt_info,
call_info=CallInfo.from_prompt_info(prompt_template.prompt_info,
start_time=time.time(),
end_time=time.time(),
api_style='batch'),
)
)
# add a line to the batch file using the completion id as the id
batch_file_data.append({
"custom_id": completion_info.completion_id,
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": prompt_template.prompt_info.model,
"messages": formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters,
}
})
# write the batch file somewhere persistent
with open("batch_file.jsonl", "w") as f:
for line in batch_file_data:
json.dump(line, f)
f.write("\n")
## Upload the batch file to the OpenAI API ##
batch_input_file = openai_client.files.create(
file=open("batch_file.jsonl", "rb"),
purpose="batch"
)
## Create the batch request ##
batch_request = openai_client.batches.create(
input_file_id=batch_input_file.id,
endpoint="/v1/chat/completions",
completion_window="24h",
metadata={
"description": "nightly job"
}
)
print(batch_request)
### NEW SERVICE POLLS FOR COMPLETION ###
## New service does the polling to check on the status of the batch request ##
batch_status = batch_request.status
while batch_status != "completed":
time.sleep(10)
batch_request = openai_client.batches.retrieve(batch_request.id)
batch_status = batch_request.status
print(batch_status)
## Read the batch file and update in freeplay ##
file_response = openai_client.files.content(batch_request.output_file_id)
# Do whatever you want with the response
# will write to a file for now
with open("batch_output.jsonl", "wb") as f:
f.write(file_response.content)
# loop over each line and update the completion in freeplay
for line in file_response.text.strip().split('\n'):
response_data = json.loads(line)
completion_id = response_data['custom_id'] # this is the completion id for the partial completion in freeplay
output = response_data['response']['body']['choices'][0]['message']['content']
print(completion_id, output)
fp_client.recordings.update(
RecordUpdatePayload(
project_id=project_id,
completion_id=completion_id,
new_messages=[
{
"role": "assistant",
"content": output,
}
],
eval_results={}
)
)
```
# OpenAI Responses API
Source: https://docs.freeplay.ai/developer-resources/recipes/openai-responses-api
Use the OpenAI Responses API with Freeplay for text, images, tools, and structured outputs.
The OpenAI [Responses API](https://platform.openai.com/docs/api-reference/responses) is OpenAI's latest API for generating completions. It supports text, images, tool calling, structured outputs, and more in a unified interface. Freeplay supports formatting prompts and recording completions with Responses API so you can use it seamlessly in your code.
## Setting up in the prompt playground
To use the Responses API with a prompt template in Freeplay:
1. Open your prompt template in the **Prompt Playground**
2. Select a compatible OpenAI model (e.g. `gpt-5.x`, `gpt-4.x`)
3. Open **Model Settings**
4. Change the **API Format** to **Responses**
Once configured, Freeplay formats the prompt for the Responses API when fetched via the SDK. This means `formatted_prompt.llm_prompt` returns the input array expected by `openai.responses.create()` instead of the chat completions message format.
## How it works
When the API Format is set to Responses API:
* **`formatted_prompt.llm_prompt`** returns the `input` array for `responses.create()`
* **`formatted_prompt.system_content`** returns the system instructions (passed as `instructions`)
* **`formatted_prompt.tool_schema`** returns tools in the Responses API format
* **`formatted_prompt.formatted_output_schema`** returns the JSON schema for structured outputs
* **`formatted_prompt.prompt_info.model_parameters`** contains model settings like `temperature`
### 1. Setup clients
Initialize Freeplay and OpenAI client SDKs.
### 2. Fetch prompt from Freeplay
Pull in the formatted prompt. Since the API Format is set to Responses API, the prompt is formatted accordingly.
### 3. Build the Responses API call
Map the formatted prompt fields to the Responses API parameters — `instructions`, `tools`, structured output `text` format, etc.
### 4. Call OpenAI Responses API
Pass the formatted input and parameters to `openai.responses.create()`.
### 5. Handle the response
The Responses API returns an `output` array. Iterate through it to handle text outputs and tool calls.
### 6. Record the interaction
Pass the messages, tool schema, and media inputs back to Freeplay for observability.
## Examples
```python Python theme={null}
import base64
import json
import os
import time
from pathlib import Path
from typing import Any, Dict, Optional
from openai import OpenAI
from freeplay import Freeplay, RecordPayload, CallInfo
from freeplay.model import MediaInputBase64
from freeplay.resources.recordings import UsageTokens
## SETUP ##
fp_client = Freeplay(
freeplay_api_key=os.environ["FREEPLAY_API_KEY"],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api",
)
openai_client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
project_id = os.environ["FREEPLAY_PROJECT_ID"]
input_variables = {"location": "San Francisco"}
## IMAGE INPUT (OPTIONAL) ##
image_url: Optional[str] = None # Set to an image file path to include an image input
media_inputs = {}
if image_url:
image_path = Path(image_url)
with open(image_path, "rb") as f:
encoded_image = base64.b64encode(f.read()).decode("utf-8")
media_inputs["image_input"] = MediaInputBase64(
type="base64",
content_type="image/jpeg",
data=encoded_image,
)
## PROMPT FETCH ##
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="my-openai-prompt",
environment="latest",
variables=input_variables,
media_inputs=media_inputs if media_inputs else None,
)
## BUILD RESPONSES API PARAMS ##
response_params: Dict[str, Any] = {
**formatted_prompt.prompt_info.model_parameters,
}
if formatted_prompt.system_content:
response_params["instructions"] = formatted_prompt.system_content
if formatted_prompt.tool_schema:
response_params["tools"] = formatted_prompt.tool_schema
if formatted_prompt.formatted_output_schema:
response_params["text"] = {
"format": {
"type": "json_schema",
"strict": True,
"schema": formatted_prompt.formatted_output_schema,
"name": "structured_output",
}
}
## LLM CALL ##
start = time.time()
completion = openai_client.responses.create(
input=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
**response_params,
)
end = time.time()
## HANDLE RESPONSE ##
messages = [*formatted_prompt.llm_prompt]
for output in completion.output:
if output.type == "function_call":
tool_name = output.name
tool_args = output.arguments
tool_id = output.id
# Replace with your actual tool implementation
tool_result = "70 and sunny"
messages = [
*messages,
{
"role": "user",
"content": str(tool_result),
"tool_call_id": tool_id,
"name": tool_name,
},
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": tool_id,
"function": {
"name": tool_name,
"arguments": json.dumps(tool_args),
},
"type": "function",
}
],
},
]
elif output.type == "output_text":
messages = [
*messages,
{"role": "assistant", "content": str(output.content[0].text)},
]
## RECORD ##
session = fp_client.sessions.create()
call_info = CallInfo.from_prompt_info(
formatted_prompt.prompt_info,
start,
end,
UsageTokens(completion.usage.input_tokens, completion.usage.output_tokens),
)
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=call_info,
tool_schema=formatted_prompt.tool_schema,
media_inputs=media_inputs if media_inputs else None,
)
)
```
```javascript Node theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
import OpenAI from "openai";
import fs from "fs";
// SETUP //
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: process.env["FREEPLAY_API_URL"],
});
const openaiClient = new OpenAI({ apiKey: process.env["OPENAI_API_KEY"] });
const projectId = process.env["FREEPLAY_PROJECT_ID"];
const inputVariables = { location: "San Francisco" };
// IMAGE INPUT (OPTIONAL) //
const imageUrl = null; // e.g. "/path/to/image.jpg"
const mediaInputs = {};
if (imageUrl) {
const encoded = fs.readFileSync(imageUrl).toString("base64");
mediaInputs["image_input"] = {
type: "base64",
contentType: "image/jpeg",
data: encoded,
};
}
// PROMPT FETCH //
const formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "my-openai-prompt",
environment: "latest",
variables: inputVariables,
...(Object.keys(mediaInputs).length > 0 ? { media: mediaInputs } : {}),
});
// BUILD RESPONSES API PARAMS //
const responseParams = {
...(formattedPrompt.promptInfo.modelParameters || {}),
};
if (formattedPrompt.systemContent) {
responseParams.instructions = formattedPrompt.systemContent;
}
if (formattedPrompt.toolSchema) {
responseParams.tools = formattedPrompt.toolSchema;
}
if (formattedPrompt.outputSchema) {
responseParams.text = {
format: {
type: "json_schema",
strict: true,
schema: formattedPrompt.outputSchema,
name: "structured_output",
},
};
}
// LLM CALL //
const startTime = new Date();
const completion = await openaiClient.responses.create({
input: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
...responseParams,
});
const endTime = new Date();
// HANDLE RESPONSE //
let messages = [...formattedPrompt.llmPrompt];
for (const output of completion.output) {
if (output.type === "function_call") {
const toolName = output.name;
const toolArgs = output.arguments;
const toolId = output.id;
// Replace with your actual tool implementation
const toolResult = "70 and sunny";
messages = [
...messages,
{
role: "user",
content: String(toolResult),
tool_call_id: toolId,
name: toolName,
},
{
role: "assistant",
content: null,
tool_calls: [
{
id: toolId,
function: {
name: toolName,
arguments: JSON.stringify(toolArgs),
},
type: "function",
},
],
},
];
} else if (output.type === "output_text") {
messages = [
...messages,
{ role: "assistant", content: String(output.content[0].text) },
];
}
}
// RECORD //
const session = fpClient.sessions.create();
const callInfo = getCallInfo(
formattedPrompt.promptInfo,
startTime,
endTime,
{
promptTokens: completion.usage.input_tokens,
completionTokens: completion.usage.output_tokens,
},
);
await fpClient.recordings.create({
projectId,
allMessages: messages,
sessionInfo: getSessionInfo(session),
inputs: inputVariables,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo,
toolSchema: formattedPrompt.toolSchema,
...(Object.keys(mediaInputs).length > 0 ? { mediaInputs } : {}),
});
```
```java Java theme={null}
import ai.freeplay.client.Freeplay;
import ai.freeplay.client.media.MediaInputCollection;
import ai.freeplay.client.resources.prompts.ChatMessage;
import ai.freeplay.client.resources.prompts.FormattedPrompt;
import ai.freeplay.client.resources.prompts.PromptInfo;
import ai.freeplay.client.resources.prompts.Prompts;
import ai.freeplay.client.resources.recordings.CallInfo;
import ai.freeplay.client.resources.recordings.RecordPayload;
import ai.freeplay.client.resources.recordings.RecordResponse;
import ai.freeplay.client.resources.sessions.Session;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ArrayNode;
import com.fasterxml.jackson.databind.node.ObjectNode;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.ArrayList;
import java.util.List;
import java.util.Map;
import static ai.freeplay.client.Freeplay.Config;
public class OpenAIResponsesApi {
private static final ObjectMapper mapper = new ObjectMapper();
private static final HttpClient http = HttpClient.newHttpClient();
/* SETUP */
static final String FREEPLAY_API_KEY = System.getenv("FREEPLAY_API_KEY");
static final String OPENAI_API_KEY = System.getenv("OPENAI_API_KEY");
static final String PROJECT_ID = System.getenv("FREEPLAY_PROJECT_ID");
static final String API_BASE = System.getenv("FREEPLAY_API_URL");
static final Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(FREEPLAY_API_KEY)
.baseUrl(API_BASE));
public static void main(String[] args) throws Exception {
Map inputVariables = Map.of("location", "San Francisco");
/* PROMPT FETCH */
FormattedPrompt> prompt = fpClient.prompts()
.getFormatted(new Prompts.GetFormattedRequest(
PROJECT_ID, "my-openai-prompt", "latest", inputVariables)
.mediaInputs(new MediaInputCollection()))
.get();
PromptInfo promptInfo = prompt.getPromptInfo();
String systemContent = prompt.getSystemContent().orElse(null);
@SuppressWarnings("unchecked")
List boundMessages = (List) prompt.getBoundMessages();
List
# Overview
Source: https://docs.freeplay.ai/developer-resources/recipes/overview
Complete, runnable code examples for common Freeplay integration patterns.
Code Recipes are self-contained examples that demonstrate how to accomplish specific tasks with Freeplay. Each recipe includes complete, working code that you can copy and adapt for your own projects.
## How Recipes Relate to Other Documentation
| Resource | Purpose | When to Use |
| ------------------------------------------- | ---------------------------------- | ------------------------------------------------------ |
| **Code Recipes** | Complete, runnable examples | Starting a new integration or looking for working code |
| [SDK Documentation](/freeplay-sdk/setup) | Detailed API reference | Understanding all available methods and options |
| [API Reference](/api-reference) | HTTP endpoint documentation | Building custom integrations or debugging |
| [How-To Guides](/practical-guides/overview) | Step-by-step implementation guides | Learning patterns and best practices |
## Recipe Categories
### Basic Patterns
Get started with fundamental Freeplay integration patterns.
* [Single Prompt](/developer-resources/recipes/single-prompt) - Fetch and call a single prompt template
* [Multi-Prompt Chain](/developer-resources/recipes/multi-prompt-chain) - Chain multiple prompts together
* [Record Agent Traces](/developer-resources/recipes/record-traces) - Group related completions into traces
* [Multi-Chain Prompt with Traces](/developer-resources/recipes/multi-chain-prompt-with-traces) - Complex multi-step workflows
### Chat & Conversations
Build conversational AI applications.
* [Managing Multi-Turn Chat History](/developer-resources/recipes/continuous-chat) - Maintain conversation context across turns
### Tool Calling
Implement function calling with different providers.
* [OpenAI Function Calls](/developer-resources/recipes/using-tools-with-openai) - Use tools with OpenAI models
* [Anthropic Tools](/developer-resources/recipes/using-tools-with-anthropic) - Use tools with Claude models
* [Google GenAI Chat with Tools](/developer-resources/recipes/using-tools-with-google-genai) - Multi-turn chat with tools using Google Gemini models
### Providers & Model Hosting
Connect to models hosted on different platforms.
* [OpenAI on Azure](/developer-resources/recipes/call-openai-on-azure) - Use Azure-hosted OpenAI models
* [Anthropic on Bedrock](/developer-resources/recipes/call-anthropic-on-bedrock) - Use Claude via AWS Bedrock
* [Llama on SageMaker](/developer-resources/recipes/call-llama-3-on-aws-sagemaker) - Deploy Llama models on SageMaker
* [Provider Switching](/developer-resources/recipes/provider-switching) - Switch between providers dynamically
* [LiteLLM Proxy](/developer-resources/recipes/provider-switching-with-litellm) - Use LiteLLM for unified provider access
### Testing
Run tests and evaluations programmatically.
* [Test Runs](/developer-resources/recipes/test-run) - Execute test runs against datasets
* [Tests with Tools](/developer-resources/recipes/run-a-test-with-tools-programmatically) - Run tests that include tool calls
### Specialized
Advanced patterns and integrations.
* [OpenAI Responses API](/developer-resources/recipes/openai-responses-api) - Use OpenAI's Responses API with text, images, tools, and structured outputs
* [Structured Outputs](/developer-resources/recipes/structured-outputs) - Get typed, structured responses from models
* [OpenAI Batch API](/developer-resources/recipes/openai-batch-api) - Process large batches efficiently
* [Pipecat Observer](/developer-resources/recipes/freeplay-pipecat-observer) - Monitor Pipecat voice pipelines
* [Pipecat Processor](/developer-resources/recipes/pipecat-processor-integration) - Process audio with Pipecat
* [LangGraph](/developer-resources/integrations/langgraph) - Comprehensive LangGraph integration guide
## Using Recipes
Each recipe follows a consistent structure:
1. **Prerequisites** - What you need before starting
2. **Setup** - Environment and client configuration
3. **Implementation** - Step-by-step code with explanations
4. **Complete Example** - Full working code you can run
Most recipes include examples in Python, with some also providing TypeScript and/or Java variants. Copy the code, update the configuration values for your project, and run.
# Pipecat Processor
Source: https://docs.freeplay.ai/developer-resources/recipes/pipecat-processor-integration
Add Freeplay observability to Pipecat voice pipelines using the Processor pattern.
### 1. Initialize the Freeplay Client
### 2. Fetch your formatted prompt
### 3. Pass the prompt to Pipecats LLM Service (Processor)
### 4. Initialize the FreeplayProcessor Service
### 5. Add freeplay\_logger to your Pipeline
### 6. FreeplayProcessor
The FreeplayProcessor is based on Pipecats LLMLogObserver however, it needs to be a service and not a logger in order to log the correct number of completions.
## Examples
```python Python theme={null}
###############################################################################
# Note: It is required to modify the processes_frame function in pipecat’s
# base_llm.py to pass along the OpenAILLMContext frame,
# this makes the handling easier in the FreeplayProcessor - process_frame
###############################################################################
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
context = None
if isinstance(frame, OpenAILLMContextFrame):
context: OpenAILLMContext = frame.context
await self.push_frame(frame, direction) # Add this line here to pass frame along
elif isinstance(frame, LLMMessagesFrame):
context = OpenAILLMContext.from_messages(frame.messages)
elif isinstance(frame, VisionImageRawFrame):
context = OpenAILLMContext()
context.add_image_frame_message(
format=frame.format, size=frame.size, image=frame.image, text=frame.text
)
elif isinstance(frame, LLMUpdateSettingsFrame):
await self._update_settings(frame.settings)
else:
await self.push_frame(frame, direction)
###############################################################################
from helpers.freeplay_frame import FreeplayProcessor
from freeplay import Freeplay, SessionInfo
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base=os.getenv("FREEPLAY_API_BASE")
)
# Get the unformatted prompt from Freeplay
# Get the unformatted prompt here and then bind it later on.
# This reduces latency in the system.
unformatted_prompt = fp_client.prompts.get(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=os.getenv("PROMPT_NAME"),
environment="latest",
)
formatted_prompt = unformatted_prompt.bind(
variables={},
history=[],
).format()
# Pass the formatted prompt to the LLM
llm = OpenAILLMService(model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema if formatted_prompt.tool_schema else None,
api_key=os.getenv("OPENAI_API_KEY"),
**formatted_prompt.prompt_info.model_parameters)
# Pass the Freeplay client to the FreeplayProcessor
freeplay_processor = FreeplayProcessor(fp_client=fp_client, template_name="voice-assistant", session=session, debug=True)
#....Additional Pipeline Configuration...
pipeline = Pipeline(
[
transport.input(), # Websocket input from client
stt, # Speech-To-Text
context_aggregator.user(),
llm, # LLM
tts, # Text-To-Speech
freeplay_processor, # Freeplay Logger (after tts so it can capture assistant audio)
transport.output(), # Websocket output to client
audiobuffer, # Used to buffer the audio in the pipeline
context_aggregator.assistant(),
]
)
#####################
# FreeplayLLMLogger #
#####################
import os
import io
import wave
import time
import base64
import datetime
from pipecat.frames.frames import (
Frame,
LLMFullResponseStartFrame,
LLMFullResponseEndFrame,
UserStartedSpeakingFrame,
UserStoppedSpeakingFrame,
InputAudioRawFrame,
BotStartedSpeakingFrame,
BotStoppedSpeakingFrame,
MetricsFrame,
TTSAudioRawFrame,
)
from pipecat.processors.aggregators.openai_llm_context import OpenAILLMContextFrame
from pipecat.processors.frame_processor import FrameProcessor, FrameDirection
from freeplay import (
Freeplay,
RecordPayload,
CallInfo,
SessionInfo,
)
from freeplay.resources.prompts import PromptInfo
from pipecat.metrics.metrics import ProcessingMetricsData, TTFBMetricsData
class FreeplayProcessor(FrameProcessor):
"""Logs LLM interactions and audio to Freeplay with simplified structure."""
def __init__(
self,
fp_client: Freeplay,
template_name: str,
session: SessionInfo = None,
required_information: str = None,
unformatted_prompt: PromptInfo = None,
):
super().__init__()
self.fp_client = fp_client
self.template_name = template_name
self.conversation_id = self._new_conv_id()
self.total_completion_time = 0
self.required_information = required_information
self.deepgram_latency = 0
# Audio related properties
self.sample_width = 2
self.sample_rate = 8000
self.num_channels = 1
self._user_audio = bytearray()
self._bot_audio = bytearray()
self.user_speaking = False
self.bot_speaking = False
# Freeplay related properties
self.conversation_history = []
self.session = session
self.most_recent_user_message = None
self.most_recent_completion = None
self.unformatted_prompt = unformatted_prompt
self.reset_recent_messages()
def _new_conv_id(self) -> str:
"""Generate a new conversation ID based on the current timestamp (this represents a customer id or similar)."""
return datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
def reset_recent_messages(self):
"""Reset all temporary message and audio storage."""
self.most_recent_user_message = None
self.most_recent_completion = None
self._user_audio = bytearray()
self._bot_audio = bytearray()
self.total_completion_time = 0
self.deepgram_latency = 0
async def process_frame(self, frame: Frame, direction: FrameDirection):
"""Process incoming frames and handle Freeplay logging."""
await super().process_frame(frame, direction)
# Handle LLM response frames
if isinstance(frame, (LLMFullResponseStartFrame, LLMFullResponseEndFrame)):
event = "START" if isinstance(frame, LLMFullResponseStartFrame) else "END"
print(f"LLMFullResponseFrame: {event}", flush=True)
# Handle LLM context frame - this is where we log to Freeplay
elif isinstance(frame, OpenAILLMContextFrame):
messages = frame.context.messages
# Extract user message and completion from context
user_messages = [m for m in messages if m.get("role") == "user"]
if user_messages:
self.most_recent_user_message = user_messages[-1].get("content")
completions = [m for m in messages if m.get("role") == "assistant"]
if completions:
self.most_recent_completion = completions[-1].get("content")
# Log to Freeplay when we have both user input and completion
if self.most_recent_user_message and self.most_recent_completion:
self._record_to_freeplay()
# Handle audio state changes
elif isinstance(frame, UserStartedSpeakingFrame):
self.user_speaking = True
elif isinstance(frame, UserStoppedSpeakingFrame):
self.user_speaking = False
elif isinstance(frame, BotStartedSpeakingFrame):
self.bot_speaking = True
elif isinstance(frame, BotStoppedSpeakingFrame):
self.bot_speaking = False
# # Handle audio data
elif isinstance(frame, InputAudioRawFrame):
if self.user_speaking:
self._user_audio.extend(frame.audio)
elif isinstance(frame, TTSAudioRawFrame):
if self.bot_speaking:
self._bot_audio.extend(frame.audio)
# Handle metrics for LLM completion time
elif isinstance(frame, MetricsFrame):
self.metrics = frame.data
for metric in frame.data:
if isinstance(metric, ProcessingMetricsData):
if "LLMService" in metric.processor:
self.total_completion_time = metric.value
elif isinstance(metric, TTFBMetricsData):
if "DeepgramSTTService" in metric.processor:
self.deepgram_latency += metric.value
# Pass frame to next processor
await self.push_frame(frame, direction)
def _record_to_freeplay(self):
"""Record the current conversation state to Freeplay."""
# Create a new trace for this interaction
trace = self.session.create_trace(
input=self.most_recent_user_message,
custom_metadata={
"deepgram_latency": self.deepgram_latency,
},
)
self.conversation_history.append(
{
"role": "user",
"content": [
{"type": "text", "text": self.most_recent_user_message},
{
"type": "input_audio",
"input_audio": {
"data": base64.b64encode(
self._make_wav_bytes(
self._user_audio, prepend_silence_secs=1
)
).decode("utf-8"),
"format": "wav",
},
},
],
},
)
# Bind the variables to the prompt
if self.unformatted_prompt:
formatted = self.unformatted_prompt.bind(
variables={"required_information": self.required_information},
history=self.conversation_history,
).format()
else:
# Get formatted prompt. Note this adds latency to the pipeline
formatted = self.fp_client.prompts.get_formatted(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=self.template_name,
environment="latest",
history=self.conversation_history,
variables={"required_information": self.required_information},
)
# Calculate latency for the LLM interaction
start, end = time.time(), time.time() + self.total_completion_time
try:
# Prepare metadata and record payload
custom_metadata = {
"conversation_id": str(self.conversation_id),
}
# Add assistant's response to conversation history
last_message = {
"role": "assistant",
"content": [
{"type": "text", "text": self.most_recent_completion},
],
"audio": {
"id": self.conversation_id,
"data": base64.b64encode(
self._make_wav_bytes(self._bot_audio, prepend_silence_secs=1)
).decode("utf-8"),
"expires_at": 1729234747,
"transcript": self.most_recent_completion,
},
}
self.conversation_history.append(last_message)
# Create recording in Freeplay
self.fp_client.recordings.create(
RecordPayload(
project_id=PROJECT_ID,
all_messages=[
*formatted.llm_prompt,
last_message, # Add the last message to the record call
],
session_info=SessionInfo(
self.session.session_id, custom_metadata=custom_metadata
),
inputs={"required_information": self.required_information},
prompt_version_info=formatted.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted.prompt_info, start, end
),
trace_info=trace,
)
)
# Record output to trace
trace.record_output(
os.getenv("FREEPLAY_PROJECT_ID"),
self.most_recent_completion,
)
print(
f"Successfully recorded to Freeplay - completion time: {self.total_completion_time}s",
flush=True,
)
self.reset_recent_messages()
except Exception as e:
print(f"Error recording to Freeplay: {e}", flush=True)
self.reset_recent_messages()
def _make_wav_bytes(self, pcm: bytes, prepend_silence_secs: int = 1) -> bytes:
"""Convert PCM audio data to WAV format with optional silence prepend."""
buf = io.BytesIO()
if prepend_silence_secs > 0:
silence_samples = int(
self.sample_rate
* self.sample_width
* self.num_channels
* prepend_silence_secs
)
silence = b"\x00" * silence_samples
pcm = silence + pcm
with wave.open(buf, "wb") as wf:
wf.setnchannels(self.num_channels)
wf.setsampwidth(self.sample_width)
wf.setframerate(self.sample_rate)
wf.writeframes(pcm)
return buf.getvalue()
```
# Provider Switching
Source: https://docs.freeplay.ai/developer-resources/recipes/provider-switching
Switch between LLM providers dynamically using Freeplay prompt configuration.
### 1. Create the necessary clients
### 2. Fetch the prompt from Freeplay
Freeplay will format your messages and model parameters for the provider you have configured on the prompt template version
### 3. Make the LLM call
Key off the provider name from the formatted prompt and route to the proper provider
### 4. Record to Freeplay
Record the interaction back to Freeplay
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo
from openai import OpenAI
from anthropic import Anthropic
import os
import time
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base=os.getenv("FREEPLAY_URL"),
)
project_id=os.getenv("FREEPLAY_PROJECT_ID")
openai_client = OpenAI(
api_key=os.getenv("OPENAI_API_KEY"),
)
anthropic_client = Anthropic(
api_key=os.getenv("ANTHROPIC_API_KEY"),
)
user_question = "What is the capital of France?"
# Fetch the prompt from freeplay
prompt_vars = {"question": user_question}
formatted_prompt = fp_client.prompts.get_formatted(
project_id="5688ebaf-7f22-4d5d-b9bb-bc715c8faabb",
template_name="BasicTriviaBot",
environment="dev",
variables=prompt_vars,
)
# make the llm call using the provider configured on the prompt
# Freeplay will format the prompt and model parameters for the given provider
start = time.time()
if formatted_prompt.prompt_info.provider == "openai":
response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
msg = response.choices[0].message
answer = msg.content
elif formatted_prompt.prompt_info.provider == "anthropic":
response = anthropic_client.messages.create(
model=formatted_prompt.prompt_info.model,
system=formatted_prompt.system_content or NotGiven(), # Freeplay splits out the system prompt from messages for Anthropic
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
msg = response
answer = msg.content[0].text
else:
raise ValueError(f"Unsupported provider: {formatted_prompt.prompt_info.provider}")
end = time.time()
# record to Freeplay
session = fp_client.sessions.create()
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
tool_schema=formatted_prompt.tool_schema,
session_info=session.session_info,
inputs=test_case.variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end)
)
)
print(f'Question {user_question} \n Answer {answer}')
```
# Provider Switching with LiteLLM
Source: https://docs.freeplay.ai/developer-resources/recipes/provider-switching-with-litellm
Route LLM calls through LiteLLM while recording to Freeplay.
### 1. Create the Freeplay Client
### 2. Fetch the prompt from Freeplay
Freeplay will handle conversion of the messages to the right format for you
### 3. Call the LLM via LiteLLM
Route the LLM call through LiteLLM
### 4. Record to Freeplay
Record the interaction back to Freeplay
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo
from openai import OpenAI
from anthropic import Anthropic
from dotenv import load_dotenv
import os
import time
from litellm import completion
load_dotenv("../.env")
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base=os.getenv("FREEPLAY_URL"),
)
user_question = "What is the capital of France?"
# Fetch the prompt from freeplay
prompt_vars = {"question": user_question}
formatted_prompt = fp_client.prompts.get_formatted(
project_id="5688ebaf-7f22-4d5d-b9bb-bc715c8faabb",
template_name="BasicTriviaBot",
environment="dev",
variables=prompt_vars,
)
start = time.time()
response = completion(
model=f"{formatted_prompt.prompt_info.provider}/{formatted_prompt.prompt_info.model}",
messages=formatted_prompt.llm_prompt,
)
msg = response.choices[0].message
answer = msg.content
end = time.time()
# record to Freeplay
session = fp_client.sessions.create()
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=formatted_prompt.all_messages(msg),
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end)
)
)
print(f'Question {user_question} \n Answer {answer}')
```
# Record Traces
Source: https://docs.freeplay.ai/developer-resources/recipes/record-traces
Record agent traces containing multiple LLM completions to Freeplay for observability.
### 1. Create client
### 2. Define a Call and Record Helper
### 3. Pass through Trace Info to Record
Pass trace info through on the record call to freeplay
### 4. Loop over questions and record to traces
### 5. Create a Trace Object
Create a Trace Object, including a user display input message
### 6. Close and Record the Trace
Record the Trace with the final display output to close the trace
## Examples
```python Python theme={null}
import os
import random
import time
from typing import Optional
from anthropic import Anthropic, NotGiven
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo, TraceInfo
fp_client = Freeplay(
freeplay_api_key=os.environ['FREEPLAY_API_KEY'],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api"
)
project_id = os.environ['FREEPLAY_PROJECT_ID']
client = Anthropic(
api_key=os.environ.get("ANTHROPIC_API_KEY")
)
def call_and_record(
project_id: str,
template_name: str,
env: str,
input_variables: dict,
session_info: SessionInfo,
trace_info: Optional[TraceInfo] = None
) -> dict:
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name=template_name,
environment=env,
variables=input_variables
)
print(f"Ready for LLM: {formatted_prompt.llm_prompt}")
start = time.time()
completion = client.messages.create(
system=formatted_prompt.system_content or NotGiven(),
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
llm_response = completion.content[0].text
print("Completion: %s" % llm_response)
all_messages = formatted_prompt.all_messages(
new_message={'role': 'assistant', 'content': llm_response}
)
call_info = CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end)
record_response = fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
session_info=session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=call_info,
trace_info=trace_info
)
)
return {'completion_id': record_response.completion_id, 'llm_response': llm_response}
# send 3 questions to the model encapsulated into a trace
user_questions = ["answer life's most existential questions", "what is sand?", "how tall are lions?"]
session = fp_client.sessions.create()
for question in user_questions:
trace_info = session.create_trace(input=question)
bot_response = call_and_record(
project_id=project_id,
template_name='my-anthropic-prompt',
env='latest',
input_variables={'question': question},
session_info=session.session_info,
trace_info=trace_info
)
categorization_result = call_and_record(
project_id=project_id,
template_name='question-classifier',
env='latest',
input_variables={'question': question},
session_info=session.session_info,
trace_info=trace_info
)
trace_info.record_output(project_id, bot_response['llm_response'])
print(f"Trace info id: {trace_info.trace_id}")
```
```javascript Node theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay/thin";
import Anthropic from "@anthropic-ai/sdk";
const projectId = process.env["FREEPLAY_PROJECT_ID"];
const environment = "latest";
const templateName = "my-prompt-anthropic";
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: `${process.env["FREEPLAY_API_URL"]}/api`,
});
const anthropicClient = new Anthropic({
apiKey: process.env["ANTHROPIC_API_KEY"],
});
async function call(
projectId,
templateName,
environment,
input_variables,
session_info,
trace_info
) {
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName,
environment,
variables: input_variables,
});
let start = new Date();
const llmResponse = await anthropicClient.messages.create({
model: formattedPrompt.promptInfo.model,
messages: formattedPrompt.llmPrompt,
system: formattedPrompt.systemContent,
...formattedPrompt.promptInfo.modelParameters,
});
let end = new Date();
const llmResponseText = llmResponse.content[0].text;
let messages = formattedPrompt.allMessages({
content: llmResponseText,
role: "Assistant",
});
const completionResponse = await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: input_variables,
sessionInfo: session_info,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
traceInfo: trace_info,
});
return {
completionId: completionResponse.completionId,
llmResponseText: llmResponseText,
};
}
const userQuestion = "answer life's most existential questions";
const session = await fpClient.sessions.create({
customMetadata: { some_custom_metadata: 42 },
});
const traceInfo = await session.createTrace(userQuestion);
const botResponse = await call(
projectId,
templateName,
environment,
{ question: userQuestion },
getSessionInfo(session),
traceInfo
);
const categorizationResponse = await call(
projectId,
templateName,
environment,
{ question: `categorize this question ${userQuestion}` },
getSessionInfo(session),
traceInfo
);
await traceInfo.recordOutput(projectId, botResponse.llmResponseText);
console.log(
`Trace recorded with Id ${traceInfo.traceId} and input "${traceInfo.input}" and output "${botResponse.llmResponseText}"`
);
```
```java Java theme={null}
package ai.freeplay.example.java;
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.ChatMessage;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.recordings.CallInfo;
import ai.freeplay.client.thin.resources.recordings.RecordInfo;
import ai.freeplay.client.thin.resources.sessions.Session;
import ai.freeplay.client.thin.resources.sessions.TraceInfo;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.net.http.HttpResponse;
import java.util.List;
import java.util.Map;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import static ai.freeplay.client.thin.Freeplay.Config;
import static ai.freeplay.example.java.ThinExampleUtils.callAnthropic;
import static java.lang.String.format;
public class ThinTrace {
static String freeplayApiKey = System.getenv("FREEPLAY_API_KEY");
static String projectId = System.getenv("FREEPLAY_PROJECT_ID");
static String customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME");
static String anthropicApiKey = System.getenv("ANTHROPIC_API_KEY");
private static final ObjectMapper objectMapper = new ObjectMapper();
static Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
);
public static String call(
String projectId,
String templateName,
String environment,
Map variables,
Session session,
TraceInfo traceInfo
) {
return fpClient.prompts()
.>getFormatted(
projectId,
templateName,
environment,
variables,
null
).thenCompose((FormattedPrompt> formattedPrompt) -> {
long startTime = System.currentTimeMillis();
return callAnthropic(
objectMapper,
anthropicApiKey,
formattedPrompt.getPromptInfo().getModel(),
formattedPrompt.getPromptInfo().getModelParameters(),
formattedPrompt.getFormattedPrompt(),
formattedPrompt.getSystemContent().orElse(null)
).thenApply((HttpResponse response) ->
new ThinExampleUtils.Tuple3<>(formattedPrompt, response, startTime)
);
}
).thenCompose((ThinExampleUtils.Tuple3>, HttpResponse, Long> promptAndResponse) -> {
FormattedPrompt> formattedPrompt = promptAndResponse.first;
HttpResponse response = promptAndResponse.second;
long startTime = promptAndResponse.third;
JsonNode bodyNode;
try {
bodyNode = objectMapper.readTree(response.body());
} catch (JsonProcessingException e) {
throw new RuntimeException("Unable to parse response body.", e);
}
List allMessages = formattedPrompt.allMessages(
new ChatMessage("assistant", bodyNode.path("content").get(0).path("text").asText())
);
CallInfo callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
System.currentTimeMillis()
);
String output = bodyNode.path("content").get(0).path("text").asText();
System.out.println("Completion: " + output);
RecordInfo recordInfo = new RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo)
.traceInfo(traceInfo));
recordInfo.traceInfo(traceInfo);
fpClient.recordings().create(recordInfo);
return CompletableFuture.completedFuture(output);
}
)
.exceptionally(exception -> {
System.out.println("Got exception: " + exception.getMessage());
return null;
})
.join();
}
public static void main(String[] args) throws ExecutionException, InterruptedException{
String input = "What is the meaning of life?";
Map inputVars = Map.of("question", input);
Session session = fpClient.sessions().create();
TraceInfo traceInfo = session.createTrace(input);
String response = call(
projectId,
"my-anthropic-prompt",
"latest",
inputVars,
session,
traceInfo
);
System.out.println("First Completion: " + response);
Map inputVars2 = Map.of("question", format("categorize the following question: %s", input));
String category = call(
projectId,
"my-anthropic-prompt",
"latest",
inputVars2,
session,
traceInfo
);
traceInfo.recordOutput(projectId, response);
System.out.println("Second Completion: " + category);
System.out.println("Recorded Trace " + traceInfo.traceId + " to session " + traceInfo.sessionId + " with input " + traceInfo.input + " and output " + response);
}
}
```
# Run a Test with Tools Programmatically
Source: https://docs.freeplay.ai/developer-resources/recipes/run-a-test-with-tools-programmatically
Run programmatic test runs with tool calling functionality using the Freeplay SDK.
### 1. Setup Freeplay and LLM SDK
Initialize Freeplay and your AI Provider's SDK (OpenAI in this example).
### 2. Fetch raw prompt
Use Freeplay SDK to pull in your raw prompt. This prompt template contains the tool schema you saved in Freeplay web app.
We are using the raw prompt to bind with test cases specific variable and history.
### 3. Create a test run
With Freeplay SDK, create a test run. We'll use this down below to associate test related data to this run.
### 4. Format prompt with test case variables
For each test case in a test run, we'll bind its variable and history with the prompt we fetched
### 5. Call LLM with the tools
Create a new completion and pass in the tool schema from formatted prompt when creating a new completion.
### 6. Capture test run details with Freeplay
Record the eval result with its associated messages with tool calls and schema.
## Examples
```python Python theme={null}
import os
import time
from openai import OpenAI
from freeplay import Freeplay, RecordPayload, CallInfo
fp_client = Freeplay(freeplay_api_key=os.environ['FREEPLAY_API_KEY'])
openai_client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
project_id = os.environ['FREEPLAY_PROJECT_ID']
template_prompt = fp_client.prompts.get(
project_id=project_id,
template_name='your-prompt',
environment='latest'
)
test_run = fp_client.test_runs.create(
project_id,
"Name of your dataset",
include_outputs=True,
name=f'My Example Test Run',
description='Run from examples',
flavor_name=template_prompt.prompt_info.flavor_name
)
for test_case in test_run.test_cases:
formatted_prompt = template_prompt.bind(test_case.variables, history=test_case.history).format()
start = time.time()
completion = openai_client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
session = fp_client.sessions.create()
all_messages = formatted_prompt.all_messages(completion.choices[0].message)
# Handle tool call and append its result to all_messages.
# Look at OpenAI Recipe: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-openai
# Anthropic: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-anthropic
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
tool_schema=formatted_prompt.tool_schema,
session_info=session.session_info,
inputs=test_case.variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
test_run_info=test_run.get_test_run_info(test_case.id),
eval_results={
'f1-score': 0.48,
'is_non_empty': True
}
)
)
```
```javascript Node theme={null}
import Freeplay, { getSessionInfo, getTestRunInfo } from "freeplay";
import OpenAI from "openai";
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: `${process.env["FREEPLAY_API_URL"]}/api`,
});
const openaiClient = new OpenAI({
apiKey: process.env["OPENAI_API_KEY"],
});
const projectId = process.env["FREEPLAY_PROJECT_ID"];
const templatePrompt = await fpClient.prompts.get({
projectId,
templateName: "your-prompt",
environment: "latest",
});
const testRun = await fpClient.testRuns.create({
projectId,
testList: "Name of your dataset",
includeOutputs: true,
name: "My Example Test Run",
description: "Run from examples",
flavorName: templatePrompt.promptInfo.flavorName,
});
for await (const testCase of testRun.testCases) {
const formattedPrompt = templatePrompt
.bind(testCase.variables, testCase.history)
.format();
const start = new Date();
const completion = await openaiClient.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
tools: formattedPrompt.toolSchema,
...formattedPrompt.promptInfo.modelParameters,
});
const end = new Date();
const session = fpClient.sessions.create();
const messages = formattedPrompt.allMessages(completion.choices[0].message);
// Handle tool call and append its result to all_messages.
// Look at OpenAI Recipe: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-openai
// Anthropic: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-anthropic
await fpClient.recordings.create({
projectId,
allMessages: messages,
toolSchema: formattedPrompt.toolSchema,
sessionInfo: getSessionInfo(session),
inputs: testCase.variables,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: {
provider: formattedPrompt.promptInfo.provider,
model: formattedPrompt.promptInfo.model,
startTime: start,
endTime: end,
modelParameters: formattedPrompt.promptInfo.modelParameters,
},
testRunInfo: getTestRunInfo(testRun, testCase.id),
evalResults: {
"f1-score": 0.48,
is_non_empty: true,
},
});
}
```
```java Kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.*
import ai.freeplay.client.thin.resources.recordings.*
import ai.freeplay.example.java.ThinExampleUtils.callOpenAIWithTools
import com.fasterxml.jackson.databind.ObjectMapper
import java.time.Instant
class TestRunToolsExample {
companion object {
private val objectMapper = ObjectMapper()
@JvmStatic
fun main(args: Array) {
// Initialize clients
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(System.getenv("FREEPLAY_API_KEY"))
.baseUrl("${System.getenv("FREEPLAY_API_URL")}/api")
)
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
// Get template prompt
val templatePrompt = fpClient.prompts()
.get(projectId, "your-prompt", "latest")
.join()
// Create test run
val testRun = fpClient.testRuns().createRequest(projectId, "Name of your dataset")
.name("My Example Test Run")
.description("Run from examples")
.includeOutputs(true)
.flavorName(templatePrompt.promptInfo.flavorName)
.build()
.let { fpClient.testRuns().create(it).join() }
// Process each test case
testRun.testCases.forEach { testCase ->
val formattedPrompt = templatePrompt
.bind(testCase.variables, testCase.history)
.format>()
val startTime = Instant.now().toEpochMilli()
val response = callOpenAIWithTools(
objectMapper,
System.getenv("OPENAI_API_KEY"),
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt,
formattedPrompt.toolSchema
).join()
val session = fpClient.sessions().create()
val bodyNode = objectMapper.readTree(response.body())
val message = objectMapper.convertValue(bodyNode["choices"][0]["message"], Object::class.java)
val allMessages = formattedPrompt.allMessages(message)
// Handle tool call and append its result to all_messages.
// Look at OpenAI Recipe: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-openai
// Anthropic: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-anthropic
fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(testCase.variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(formattedPrompt.promptInfo)
.callInfo( CallInfo.from(
formattedPrompt.promptInfo,
startTime,
System.currentTimeMillis()
),)
.toolSchema(formattedPrompt.toolSchema)
.testRunInfo(testRun.getTestRunInfo(testCase.testCaseId))
.evalResults(mapOf(
"f1-score" to 0.48,
"is_non_empty" to true
))
).join()
}
}
}
}
```
# Single Prompt
Source: https://docs.freeplay.ai/developer-resources/recipes/single-prompt
Fetch a prompt from Freeplay, call an LLM, and record the completion.
### 1. Configure Freeplay Client
You'll need to set the following environment variables:
1. FREEPLAY\_API\_KEY
2. OPENAI\_API\_KEY
3. FREEPLAY\_PROJECT\_ID
### 2. Fetch Prompt
Fetch your prompt from the Freeplay server
### 3. Call your LLM
You will interact with your LLM directly, but you can key model config and messages off of the formatted prompt object
### 4. Record the LLM Interaction
Record your LLM response back to Freeplay
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo
from openai import OpenAI
# create your a freeplay client object
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api"
)
# configure your openai client
openai_client = OpenAI(
api_key=userdata.get('OPENAI_API_KEY'),
)
project_id=os.getenv("FREEPLAY_PROJECT_ID")
## PROMPT FETCH ##
# set the prompt variables
prompt_vars = {"keyA": "valueA"}
# get a formatted prompt
formatted_prompt = fp_client.prompts.get_formatted(project_id=project_id,
template_name="template_name",
environment="latest",
variables=prompt_vars)
## LLM CALL ##
# Make an LLM call to your provider of choice
start = time.time()
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# add the response to your message set
all_messages = formatted_prompt.all_messages(
{'role': chat_response.choices[0].message.role,
'content': chat_response.choices[0].message.content}
)
## RECORD ##
# create a session
session = fp_client.sessions.create()
# build the record payload
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=start, end_time=end)
)
# record the LLM interaction
fp_client.recordings.create(payload)
```
```javascript Node theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
import OpenAI from "openai";
// create your freeplay client
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: "https://acme.freeplay.ai/api",
});
/* PROMPT FETCH */
// set the prompt variables
let promptVars = {"keyA": "valueA"};
// fetch a formatted prompt template
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId: projectID,
templateName: "template_name",
environment: "latest",
variables: promptVars,
});
/* LLM CALL */
// make the llm call
const openai = new OpenAI(process.env["OPENAI_API_KEY"]);
let start = new Date();
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
});
let end = new Date();
console.log(chatCompletion.choices[0].message);
// update the messages
let messages = formattedPrompt.allMessages({
role: chatCompletion.choices[0].message.role,
content: chatCompletion.choices[0].message.content,
});
/* RECORD */
// create a session
let session = fpClient.sessions.create({});
// record the LLM interaction with Freeplay
await fpClient.recordings.create({
allMessages: messages,
inputs: promptVars,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end)
});
```
```java Kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
import ai.freeplay.example.java.ThinExampleUtils.callOpenAI
import com.fasterxml.jackson.databind.ObjectMapper
import kotlinx.coroutines.future.await
import kotlinx.coroutines.runBlocking
private val objectMapper = ObjectMapper()
fun main(): Unit = runBlocking {
// set your environment variables
val openaiApiKey = System.getenv("OPENAI_API_KEY")
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME")
// create your url from your customer domain
val baseUrl = String.format("https://%s.freeplay.ai/api", customerDomain)
// create the freeplay client
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
)
// set the prompt variables
val promptVars = mapOf("keyA" to "valueA")
/* PROMPT FETCH */
val formattedPrompt = fpClient.prompts().getFormatted(
projectId,
"template-name",
"latest",
promptVars
).await()
/* LLM CALL */
// set timer to measure latency
val startTime = System.currentTimeMillis()
val llmResponse = callOpenAI(
objectMapper,
openaiApiKey,
formattedPrompt.promptInfo.model, // get the model name
formattedPrompt.promptInfo.modelParameters, // get the model params
formattedPrompt.getBoundMessages()
).await()
val endTime = System.currentTimeMillis()
// add the LLM response to the message set
val bodyNode = objectMapper.readTree(llmResponse.body())
val role = bodyNode.path("choices").path(0).path("message").path("role").asText()
val content = bodyNode.path("choices").path(0).path("message").path("content").asText()
println("Completion: " + content)
val allMessage: List = formattedPrompt.allMessages(
ChatMessage(role, content)
)
/* RECORD */
// construct the call info from the prompt object
val callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
endTime
)
// create the session and get the session info
val sessionInfo = fpClient.sessions().create().sessionInfo
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
allMessage,
promptVars,
sessionInfo,
formattedPrompt.getPromptInfo(),
callInfo
)
).await()
println("Completion Record Succeeded with ${recordResponse.completionId}")
}
```
# Streaming Responses
Source: https://docs.freeplay.ai/developer-resources/recipes/streaming-responses
Handle streaming LLM responses while recording completions to Freeplay for observability.
### Run a single prompt and stream the result
Executes a single prompt using specified variables, with a streaming response.
## Examples
```python python theme={null}
import os
from freeplay import Freeplay
from freeplay.provider_config import ProviderConfig, OpenAIConfig
FREEPLAY_API_KEY = os.environ["FREEPLAY_API_KEY"]
OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]
FREEPLAY_CUSTOMER_NAME = os.environ["FREEPLAY_CUSTOMER_NAME"]
FREEPLAY_PROJECT_ID = os.environ["FREEPLAY_PROJECT_ID"]
fp_client = Freeplay(
provider_config=ProviderConfig(openai=OpenAIConfig(OPENAI_API_KEY)),
freeplay_api_key=FREEPLAY_API_KEY,
api_base=f'https://{FREEPLAY_CUSTOMER_NAME}.freeplay.ai/api'
)
completion_stream = fp_client.get_completion_stream(
project_id=FREEPLAY_PROJECT_ID,
template_name="album_bot",
variables={"pop_star": "Bruno Mars"}
)
for chunk in completion_stream:
print("Chunk: %s" % chunk.text.strip())
```
```javascript node theme={null}
import * as freeplay from "freeplay";
async function main() {
const FREEPLAY_API_KEY = process.env["FREEPLAY_API_KEY"]
const OPENAI_API_KEY = process.env["OPENAI_API_KEY"]
const FREEPLAY_CUSTOMER_NAME = process.env["FREEPLAY_CUSTOMER_NAME"]
const FREEPLAY_PROJECT_ID = process.env["FREEPLAY_PROJECT_ID"]
const fp_client = new freeplay.Freeplay({
freeplayApiKey: FREEPLAY_API_KEY,
baseUrl: `https://${FREEPLAY_CUSTOMER_NAME}.freeplay.ai/api`,
providerConfig: {
openai: {
apiKey: OPENAI_API_KEY
}
}
})
const completionStream = await fp_client.getCompletionStream({
projectId: FREEPLAY_PROJECT_ID,
templateName: "album_bot",
variables: {
["pop_star"]: "Bruno Mars"
}
})
for await (const chunk of completionStream) {
if (chunk.choices[0].role) {
process.stdout.write(`${chunk.choices[0].role} message: `);
}
process.stdout.write(chunk.content);
}
console.log('\n')
}
main();
```
```kotlin kotlin theme={null}
package ai.freeplay.example.java;
import ai.freeplay.client.Freeplay;
import ai.freeplay.client.ProviderConfig.OpenAIProviderConfig;
import ai.freeplay.client.model.ChatMessage;
import ai.freeplay.client.model.CompletionResponse;
import ai.freeplay.client.model.CompletionSession;
import java.util.Collections;
import java.util.Map;
import java.util.stream.Stream;
import static java.lang.String.format;
public class StreamingExample {
public static void main(String[] args) {
String openaiApiKey = System.getenv("OPENAI_API_KEY");
String freeplayApiKey = System.getenv("FREEPLAY_API_KEY");
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
String customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME");
String baseUrl = format("https://%s.freeplay.ai/api", customerDomain);
Freeplay fpClient = new Freeplay(freeplayApiKey, baseUrl, new OpenAIProviderConfig(openaiApiKey));
Map llmParameters = Collections.emptyMap();
CompletionSession session = fpClient.createSession(projectId, "prod");
Stream completionStream = session.getCompletionStream(
"my-chat-start",
Map.of("question", "why isn't my sink working?"),
llmParameters,
null,
null
);
completionStream.forEach((ChatMessage chunk) -> {
System.out.printf("Message [%s]: %s%n", chunk.getRole(), chunk.getContent());
});
Stream textStream = session.getCompletionStream(
"my-prompt",
Map.of("question", "why isn't my sink working?"),
llmParameters,
null,
null
);
textStream.forEach((CompletionResponse chunk) -> {
System.out.printf("Message [TEXT]: %s%n", chunk.getContent());
});
}
}
```
# Structured Outputs
Source: https://docs.freeplay.ai/developer-resources/recipes/structured-outputs
Use structured outputs to get typed responses from LLMs with Freeplay.
### 1. Structred outputs
This shows how to use structured outputs with Freeplay's SDK
## Examples
```python Python theme={null}
import os
import time
from typing import List
import pydantic
from openai import OpenAI
from freeplay import Freeplay, RecordPayload, CallInfo
from freeplay.resources.recordings import UsageTokens
# Define structured output classes with Pydantic
class COTStep(pydantic.BaseModel):
thinking: str
result: str
class COTResponse(pydantic.BaseModel):
response: str
steps: List[COTStep]
# Initialize clients
fp_client = Freeplay(
freeplay_api_key=os.environ["FREEPLAY_API_KEY"],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api",
)
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
input_variables = {"question": "why is the sky blue?"}
project_id = os.environ["FREEPLAY_PROJECT_ID"]
# Fetch formatted prompt with output schema
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="my-chat-template",
environment="latest",
variables=input_variables,
)
print(f"Tool schema: {formatted_prompt.tool_schema}")
print(f"Output schema: {formatted_prompt.formatted_output_schema}")
start = time.time()
# Build the completion parameters
completion_params = {
**formatted_prompt.prompt_info.model_parameters,
}
# Add tools if present
if formatted_prompt.tool_schema:
completion_params["tools"] = formatted_prompt.tool_schema
# Use the output schema from the prompt template
completion = client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
response_format={
"type": "json_schema",
"json_schema": {
"strict": True,
"schema": COTResponse.model_json_schema(), # OR you can use formatted_prompt.formatted_output_schema,
"name": "COTReasoning",
},
}
if formatted_prompt.formatted_output_schema
else openai.NotGiven(),
)
end = time.time()
print("Completion with prompt schema: %s" % completion)
# Record to Freeplay
session = fp_client.sessions.create()
messages = formatted_prompt.all_messages(completion.choices[0].message)
call_info = CallInfo.from_prompt_info(
formatted_prompt.prompt_info,
start,
end,
UsageTokens(completion.usage.prompt_tokens, completion.usage.completion_tokens),
api_style="batch",
)
record_response = fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=call_info,
tool_schema=formatted_prompt.tool_schema,
output_schema=COTResponse.model_json_schema() # OR you can use formatted_prompt.formatted_output_schema
)
)
```
```javascript JavaScript theme={null}
import OpenAI from "openai";
import { z } from "zod";
import Freeplay, { getSessionInfo } from "@freeplay/freeplay";
// Define structured output schemas using Zod
const COTStepSchema = z.object({
thinking: z.string(),
result: z.string(),
});
const COTResponseSchema = z.object({
response: z.string(),
steps: z.array(COTStepSchema),
});
async function main() {
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: `${process.env["FREEPLAY_API_URL"]}/api`,
});
const openaiClient = new OpenAI({
apiKey: process.env["OPENAI_API_KEY"],
});
const inputVariables = { question: "why is the sky blue?" };
const projectId = process.env["FREEPLAY_PROJECT_ID"];
// Fetch formatted prompt with output schema
const formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "my-chat-template",
environment: "latest",
variables: inputVariables,
});
console.log("Tool schema:", formattedPrompt.toolSchema);
console.log("Output schema:", formattedPrompt.outputSchema);
const start = new Date();
// Build the completion parameters
const completionParams: OpenAI.ChatCompletionCreateParams = {
messages: (formattedPrompt.llmPrompt || []) as OpenAI.ChatCompletionMessageParam[],
model: formattedPrompt.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters,
};
// Add tools if present
if (formattedPrompt.toolSchema) {
completionParams.tools = formattedPrompt.toolSchema;
}
// Use output schema from prompt template, or fall back to Zod schema
let completion: OpenAI.ChatCompletion;
let jsonSchema: any = undefined;
if (formattedPrompt.outputSchema) {
// Use schema from prompt template
completion = await openaiClient.chat.completions.create({
...completionParams,
response_format: {
type: "json_schema",
json_schema: {
strict: true,
schema: formattedPrompt.outputSchema,
name: "COTReasoning",
},
},
});
console.log("Completion with prompt schema:", completion);
} else {
// Alternatively, use a Zod schema directly with Zod 4's native toJSONSchema()
jsonSchema = z.toJSONSchema(COTResponseSchema);
completion = await openaiClient.chat.completions.create({
...completionParams,
response_format: {
type: "json_schema",
json_schema: {
strict: true,
schema: jsonSchema,
name: "COTReasoning",
},
},
});
console.log("Completion with Zod schema:", completion);
}
const end = new Date();
// Record to Freeplay
const session = fpClient.sessions.create();
const messages = formattedPrompt.allMessages(completion.choices[0].message);
await fpClient.recordings.create({
projectId,
allMessages: messages,
sessionInfo: getSessionInfo(session),
inputs: inputVariables,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: {
provider: formattedPrompt.promptInfo.provider,
model: formattedPrompt.promptInfo.model,
startTime: start,
endTime: end,
modelParameters: formattedPrompt.promptInfo.modelParameters,
usage: completion.usage
? {
promptTokens: completion.usage.prompt_tokens,
completionTokens: completion.usage.completion_tokens,
}
: undefined,
},
toolSchema: formattedPrompt.toolSchema,
outputSchema: formattedPrompt.outputSchema || jsonSchema
});
console.log("Recording created successfully");
}
main().catch(console.error);
```
# Test Run
Source: https://docs.freeplay.ai/developer-resources/recipes/test-run
Execute batch test runs over datasets using the Freeplay SDK.
### 1. Configure Freeplay Client
### 2. Create a Test Run
Instantiate your Test Run which will fetch the Test Cases and create a new Test Run ID
### 3. Fetch Prompt
Fetch your prompt template. Don't format it yet, we will bind and format for each Test Case
### 4. Loop over Test Cases
Loop over each Test Case, make a completion and record to Freeplay
## Examples
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, TestRunInfo
from openai import OpenAI
# create your a freeplay client object
fpClient = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://acme.freeplay.ai/api"
)
# create a new test run
test_run = fpClient.test_runs.create(project_id=project_id, testlist="test-list-name")
# get the prompt associated with the test run
template_prompt = fpClient.prompts.get(project_id=project_id,
template_name="template-name",
environment="latest"
)
# iterate over each test case
for test_case in test_run.test_cases:
# format the prompt with the test case variables
formatted_prompt = template_prompt.bind(test_case.variables).format()
# make your llm call
s = time.time()
openai_client = OpenAI(api_key=openai_key)
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
e = time.time()
# append the results to the messages
all_messages = formatted_prompt.all_messages({
'role': chat_response.choices[0].message.role,
'content': chat_response.choices[0].message.content
})
# create a session which will create a UID
session = fp_client.sessions.create()
# build the record payload
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=test_case.variables, # the variables from the test case are the inputs
session_info=session, # use the session object created above
test_run_info=test_run.get_test_run_info(test_case.id), # link the record call to the test run and test case
prompt_version_info=formatted_prompt.prompt_info, # log the prompt information
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=s, end_time=e) # log call information
)
# record the results to freeplay
fpClient.recordings.create(payload)
```
```javascript Node theme={null}
import Freeplay, { getSessionInfo, getCallInfo, getTestRunInfo} from "freeplay";
// create your freeplay client
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: "https://acme.freeplay.ai/api",
});
// create a test run
const testRun = await fpClient.testRuns.create({
projectId: fpProjectId,
testList: 'test-list-name'
});
// fetch the prompt template for the test run
let templatePrompt = await fpClient.prompts.get({
projectId: fpProjectId,
templateName: "template-name",
environment: "latest",
});
for (const testCase of testRun.testCases) {
// create a formatted prompt from the test case
const formattedPrompt = templatePrompt.bind(testCase.variables).format();
// make the llm call
let start = new Date();
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters
});
let end = new Date();
console.log(chatCompletion.choices[0].message);
// update the messages
let messages = formattedPrompt.allMessages({
role: chatCompletion.choices[0].message.role,
content: chatCompletion.choices[0].message.content,
});
// create a session
let session = fpClient.sessions.create({});
// record the test case interaction with Freeplay
await fpClient.recordings.create({
projectId: fpProjectId,
allMessages: messages,
inputs: testCase.variables,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
testRunInfo: getTestRunInfo(testRun, testCase.id)
});
}
```
```java Kotlin theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.ChatMessage;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.prompts.TemplatePrompt;
import ai.freeplay.client.thin.resources.recordings.CallInfo;
import ai.freeplay.client.thin.resources.recordings.RecordInfo;
import ai.freeplay.client.thin.resources.recordings.RecordResponse;
import ai.freeplay.client.thin.resources.sessions.SessionInfo;
import ai.freeplay.client.thin.resources.testruns.TestCase;
import ai.freeplay.client.thin.resources.testruns.TestRun;
// create the freeplay client
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
)
val testRun = fpClient.testRuns().create(projectId, "test-list").await()
val templatePrompt = fpClient.prompts().get(projectId, "template-name", "prod").await()
for (testCase in testRun.testCases) {
// format the prompt with test case variables
val formattedPrompt = templatePrompt.bind(testCase.variables).format()
// make your llm call
val startTime = System.currentTimeMillis()
val llmResponse = callOpenAI(
objectMapper,
anthropicApiKey,
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt
).await()
val bodyNode = objectMapper.readTree(llmResponse.body())
println("Recording the result")
// append the results to your message set
val allMessages = formattedPrompt.allMessages(
ChatMessage("Assistant", bodyNode.path("completion").asText())
)
val callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
System.currentTimeMillis()
)
// create a session
val sessionInfo = fpClient.sessions().create().sessionInfo
// record the test case results
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace)
.testRunInfo(testRun.getTestRunInfo(testCase.testCaseId))
).await()
println("Recorded with completionId ${recordResponse.completionId}")
}
```
# Anthropic Tools
Source: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-anthropic
Implement tool calling with Anthropic models and record to Freeplay.
### 1. Setup clients
Initialize Freeplay and Anthropic client SDKs.
### 2. Fetch prompt from Freeplay
Pull in the formatted prompt with Freeplay. The prompt contains tools schema that'll we'll pass down below.
### 3. Call Anthropic with the tools
When creating a new completion, pass in the tools schema from the prompt we fetched.
### 4. Handle tool call
When LLM responds back with a tool call, call the external function in your service. As an example here, we are calling `get_temperature` function
### 5. Record tool call and schema
Pass in the schema and completion response to capture the tool call and its associated schema.
## Examples
```python Python theme={null}
import os
import time
import json
from anthropic import Anthropic, NotGiven
from anthropic.types import ToolUseBlock
from freeplay import Freeplay, RecordPayload, CallInfo
# A mock function that gets temperature for a location
def get_temperature(location: str) -> float:
return 72.5
project_id=os.environ['FREEPLAY_PROJECT_ID']
fp_client = Freeplay(
freeplay_api_key=os.environ['FREEPLAY_API_KEY'],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api"
)
client = Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
input_variables = {'location': "Boulder, CO"}
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name='my-anthropic-prompt',
environment='latest',
variables=input_variables
)
start = time.time()
completion = client.messages.create(
system=formatted_prompt.system_content or NotGiven(),
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# Get all messages including the completion
messages = formatted_prompt.all_messages({
'content': completion.content,
'role': completion.role,
})
# Handle tool calls if present
if isinstance(completion.content, list):
for block in completion.content:
if isinstance(block, ToolUseBlock) and block.name == "weather_of_location":
temperature = get_temperature(block.input["location"])
tool_response_message = {
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": block.id,
"content": str(temperature),
}
]
}
messages.append(tool_response_message)
session = fp_client.sessions.create()
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
tool_schema=formatted_prompt.tool_schema
)
)
```
```javascript Node theme={null}
import Anthropic from "@anthropic-ai/sdk";
import Freeplay, { getSessionInfo, getCallInfo } from "freeplay";
// A mock function that gets temperature for a location
function getTemperature(location) {
return 72.5;
}
const fp_client = new Freeplay({
freeplayApiKey: process.env.FREEPLAY_API_KEY,
baseUrl: `${process.env.FREEPLAY_API_URL}/api`,
});
const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });
const inputVariables = { location: "Boulder, CO" };
const formattedPrompt = await fp_client.prompts.getFormatted({
projectId: process.env.FREEPLAY_PROJECT_ID,
templateName: "my-anthropic-prompt",
environment: "latest",
variables: inputVariables,
});
const start = new Date();
const completion = await client.messages.create({
system: formattedPrompt.systemContent || undefined,
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
tools: formattedPrompt.toolSchema,
...formattedPrompt.promptInfo.modelParameters,
});
const end = new Date();
// Get all messages including the completion
const messages = formattedPrompt.allMessages({
content: completion.content,
role: completion.role,
});
// Handle tool calls if present
if (Array.isArray(completion.content)) {
for (const block of completion.content) {
if (block.type === "tool_use" && block.name === "weather_of_location") {
const temperature = getTemperature(block.input.location);
const toolResponseMessage = {
role: "user",
content: [
{
type: "tool_result",
tool_use_id: block.id,
content: temperature.toString(),
},
],
};
messages.push(toolResponseMessage);
}
}
}
const session = fp_client.sessions.create();
await fp_client.recordings.create({
allMessages: messages,
sessionInfo: getSessionInfo(session),
inputs: inputVariables,
promptInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
toolSchema: formattedPrompt.toolSchema
});
```
```java Kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.*
import com.fasterxml.jackson.databind.ObjectMapper
import ai.freeplay.example.java.ThinExampleUtils.callAnthropic
object AnthropicToolsExample {
private val objectMapper = ObjectMapper()
// Mock weather function
private fun getTemperature(location: String): Double = 72.5
@JvmStatic
fun main(args: Array) {
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val apiRoot = System.getenv("FREEPLAY_API_URL")
val baseUrl = "${apiRoot}/api"
val anthropicApiKey = System.getenv("ANTHROPIC_API_KEY")
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.baseUrl(baseUrl)
)
val variables = mapOf("location" to "Boulder, CO")
fpClient.prompts()
.getFormatted>(
projectId,
"my-anthropic-prompt",
"latest",
variables,
null
).thenCompose { formattedPrompt ->
val startTime = System.currentTimeMillis()
callAnthropic(
objectMapper,
anthropicApiKey,
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt,
formattedPrompt.systemContent.orElse(null),
formattedPrompt.toolSchema
).thenApply { Triple(formattedPrompt, it, startTime) }
}.thenCompose { (formattedPrompt, response, startTime) ->
try {
val bodyNode = objectMapper.readTree(response.body())
val contentNode = bodyNode.get("content")
// Create initial message list
val allMessages = formattedPrompt.allMessages(
ChatMessage("assistant", objectMapper.convertValue(contentNode, List::class.java))
).toMutableList()
// Handle tool calls
if (contentNode.isArray) {
contentNode.forEach { block ->
if (block["type"]?.asText() == "tool_use" &&
block["name"]?.asText() == "weather_of_location") {
val input = block["input"]
val temperature = getTemperature(input["location"].asText())
val toolResponse = mapOf(
"role" to "user",
"content" to listOf(
mapOf(
"type" to "tool_result",
"content" to temperature.toString(),
"tool_use_id" to block["id"].asText()
)
)
)
allMessages.add(ChatMessage(toolResponse))
}
}
}
val callInfo = CallInfo.from(
formattedPrompt.promptInfo,
startTime,
System.currentTimeMillis()
)
val sessionInfo = fpClient.sessions().create().sessionInfo
// Print the completion for debugging
if (contentNode.isArray && contentNode.size() > 0) {
contentNode[0]?.get("text")?.asText()?.let { text ->
println("Completion: $text")
}
}
val record = RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace)
.toolSchema(formattedPrompt.toolSchema)
)
fpClient.recordings().create(record)
} catch (e: Exception) {
throw RuntimeException("Failed to process JSON response", e)
}
}
.exceptionally { exception ->
System.err.println("Error: ${exception.message}")
exception.printStackTrace()
RecordResponse(null)
}
.join()
}
}
```
# Google GenAI Chat with Tools
Source: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-google-genai
Implement multi-turn chat with tool calling using Google GenAI models and record to Freeplay.
### 1. Setup clients
Initialize Freeplay and Google GenAI client SDKs.
### 2. Manage chat history
Maintain a running `messages` list across turns. Pass previous messages as `history` when fetching the formatted prompt so the model has full conversation context.
### 3. Fetch prompt from Freeplay
Pull in the formatted prompt with Freeplay. The prompt contains the system instruction and model parameters.
### 4. Call Google GenAI with the tools
When creating a new completion, pass in the tools configuration and formatted prompt contents.
### 5. Handle tool call
When the model responds with a function call, execute the external function in your service. As an example here, we are calling `get_current_temperature` function. Append both the function call and the function response to the messages list to maintain history.
### 6. Record to Freeplay
Pass in the completion response and messages to capture the tool call, its result, and the full conversation history.
## Examples
```python Python theme={null}
import os
import time
from google.genai import types
from google import genai
from freeplay import Freeplay, RecordPayload, PromptInfo, SessionInfo, CallInfo
client = genai.Client(api_key=os.getenv("GOOGLE_API_KEY"))
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api",
)
prompt_template_name = "genai_tools"
project_id = os.getenv("FREEPLAY_PROJECT_ID")
EXIT_WORDS = {"exit", "quit", "bye", "goodbye"}
weather_function = {
"name": "get_current_temperature",
"description": "Gets the current temperature for a given location.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city name, e.g. San Francisco",
},
},
"required": ["location"],
},
}
tools = types.Tool(function_declarations=[weather_function])
messages = []
def execute_tool(name: str, args: dict) -> dict:
if name == "get_current_temperature":
return {
"location": args.get("location", "unknown"),
"temperature_f": 72,
"condition": "sunny",
}
return {"error": "unknown function"}
def record_response(
inputs: dict = {},
prompt_info: PromptInfo = None,
parent_id: str = None,
session_info: SessionInfo = None,
messages: list = None,
call_info: CallInfo = None,
):
return fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
inputs=inputs or {},
prompt_version_info=prompt_info,
parent_id=parent_id,
session_info=session_info,
call_info=call_info,
)
)
print("Chat started. Type 'exit' or 'quit' to leave.\n")
session = fp_client.sessions.create()
while True:
try:
question = input("You: ").strip()
except (EOFError, KeyboardInterrupt):
print("\nExiting.")
break
if not question:
continue
if question.lower() in EXIT_WORDS:
print("Goodbye!")
break
trace = session.create_trace(input=question, agent_name="Gemini Tools")
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name=prompt_template_name,
environment="latest",
variables={"user_question": question},
history=messages,
)
messages.append({"role": "user", "parts": [{"text": question}]})
contents = list(formatted_prompt.llm_prompt)
config = types.GenerateContentConfig(
system_instruction=formatted_prompt.system_content or "",
tools=[tools],
**formatted_prompt.prompt_info.model_parameters,
)
start = time.time()
response = client.models.generate_content(
model=formatted_prompt.prompt_info.model, contents=contents, config=config
)
end = time.time()
while response.candidates[0].content.parts[0].function_call:
function_call = response.candidates[0].content.parts[0].function_call
fc_args = dict(function_call.args)
print(f"\n[Tool call] {function_call.name}({fc_args})")
# Add the function call to the messages
messages.append(
{
"role": "model",
"parts": [
{"functionCall": {"name": function_call.name, "args": fc_args}}
],
}
)
# Record the function call
record_response(
inputs=fc_args,
prompt_info=formatted_prompt.prompt_info,
parent_id=trace.trace_id,
session_info=session,
messages=messages,
call_info=CallInfo.from_prompt_info(
formatted_prompt.prompt_info, start, end
),
)
# Execute the tool
result = execute_tool(function_call.name, fc_args)
print(f"[Tool result] {result}")
# Add the function response to the messages
messages.append(
{
"role": "user",
"parts": [
{
"functionResponse": {
"name": function_call.name,
"response": result,
}
}
],
}
)
contents.append(response.candidates[0].content)
contents.append(
types.Content(
parts=[
types.Part.from_function_response(
name=function_call.name, response=result
)
],
role="user",
)
)
start = time.time()
response = client.models.generate_content(
model=formatted_prompt.prompt_info.model, contents=contents, config=config
)
end = time.time()
assistant_text = response.text
print(f"\nAssistant: {assistant_text}\n")
messages.append({"role": "model", "parts": [{"text": assistant_text}]})
record_response(
inputs={"question": question},
prompt_info=formatted_prompt.prompt_info,
parent_id=trace.trace_id,
session_info=session.session_info,
messages=messages,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
)
trace.record_output(project_id=project_id, output=assistant_text)
if any(word in assistant_text.lower() for word in EXIT_WORDS):
print("The assistant ended the conversation.")
break
```
# OpenAI Function Calls
Source: https://docs.freeplay.ai/developer-resources/recipes/using-tools-with-openai
Implement function calling with OpenAI and record tool interactions to Freeplay.
### 1. Setup clients
Initialize Freeplay and OpenAI client SDKs.
### 2. Fetch prompt from Freeplay
Pull in the formatted prompt with Freeplay. The prompt contains tools schema that'll we'll pass down below.
### 3. Call OpenAI with the tools
When creating a new completion, pass in the tools schema from the prompt we fetched.
### 4. Handle tool call
When LLM responds back with a tool call, call the external function in your service. As an example here, we are calling `get_temperature` function
### 5. Record tool call and schema
Pass in the schema and completion response to capture the tool call and its associated schema.
## Examples
```python Python theme={null}
import os
import time
import json
from openai import OpenAI
from freeplay import Freeplay, RecordPayload, CallInfo
# A mock function that gets temperature for a location
def get_temperature(location: str) -> float:
return 72.5
fp_client = Freeplay(
freeplay_api_key=os.environ['FREEPLAY_API_KEY'],
api_base=f"{os.environ['FREEPLAY_API_URL']}/api"
)
client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))
project_id=os.environ['FREEPLAY_PROJECT_ID']
input_variables = {'location': "Boulder, CO"}
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name='my-openai-prompt',
environment='latest',
variables=input_variables
)
start = time.time()
completion = client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# Append the completion to list of messages
messages = formatted_prompt.all_messages(completion.choices[0].message)
if completion.choices[0].message.tool_calls:
for tool_call in completion.choices[0].message.tool_calls:
if tool_call.function.name == "weather_of_location":
args = json.loads(tool_call.function.arguments)
temperature = get_temperature(args["location"])
tool_response_message = {
"tool_call_id": tool_call.id,
"role": "tool",
"content": str(temperature),
}
messages.append(tool_response_message)
session = fp_client.sessions.create()
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
tool_schema=formatted_prompt.tool_schema
)
)
```
```javascript Node theme={null}
import OpenAI from "openai";
import Freeplay, { getSessionInfo, getCallInfo } from "freeplay";
// A mock function that gets temperature for a location
function getTemperature(location) {
return 72.5;
}
const projectId = process.env.FREEPLAY_PROJECT_ID;
const fp_client = new Freeplay({
freeplayApiKey: process.env.FREEPLAY_API_KEY,
baseUrl: `${process.env.FREEPLAY_API_URL}/api`,
});
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const inputVariables = { location: "Boulder, CO" };
const formattedPrompt = await fp_client.prompts.getFormatted({
projectId,
templateName: "my-openai-prompt",
environment: "latest",
variables: inputVariables,
});
const start = new Date();
const completion = await client.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
tools: formattedPrompt.toolSchema,
...formattedPrompt.promptInfo.modelParameters,
});
const end = new Date();
// Append the completion to list of messages
const messages = formattedPrompt.allMessages(completion.choices[0].message);
if (completion.choices[0].message.tool_calls) {
for (const toolCall of completion.choices[0].message.tool_calls) {
if (toolCall.function.name === "weather_of_location") {
const args = JSON.parse(toolCall.function.arguments);
const temperature = getTemperature(args.location);
const toolResponseMessage = {
tool_call_id: toolCall.id,
role: "tool",
content: temperature.toString(),
};
messages.push(toolResponseMessage);
}
}
}
const session = fp_client.sessions.create();
await fp_client.recordings.create({
projectId,
allMessages: messages,
sessionInfo: getSessionInfo(session),
inputs: inputVariables,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
toolSchema: formattedPrompt.toolSchema
});
```
```java Kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
import com.fasterxml.jackson.databind.ObjectMapper
import ai.freeplay.example.java.ThinExampleUtils.callOpenAIWithTools
object OpenAIToolsExample {
private val objectMapper = ObjectMapper()
// Mock weather function
private fun getTemperature(location: String): Double = 72.5
@JvmStatic
fun main(args: Array) {
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val apiRoot = System.getenv("FREEPLAY_API_URL")
val baseUrl = "${apiRoot}/api"
val openaiApiKey = System.getenv("OPENAI_API_KEY")
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.baseUrl(baseUrl)
)
val variables = mapOf("location" to "Boulder, CO")
fpClient.prompts()
.getFormatted>(
projectId,
"my-openai-prompt",
"latest",
variables,
null
).thenCompose { formattedPrompt ->
val startTime = System.currentTimeMillis()
callOpenAIWithTools(
objectMapper,
openaiApiKey,
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt,
formattedPrompt.toolSchema
).thenApply { response ->
Triple(formattedPrompt, response, startTime)
}
}.thenCompose { (formattedPrompt, response, startTime) ->
try {
val bodyNode = objectMapper.readTree(response.body())
val choicesNode = bodyNode["choices"]
val messageNode = choicesNode[0]["message"]
val message = objectMapper.convertValue(messageNode, Object::class.java)
val allMessages = formattedPrompt.allMessages(message).toMutableList()
// Handle tool calls
val toolCalls = messageNode["tool_calls"]
if (toolCalls != null && toolCalls.isArray) {
toolCalls.forEach { toolCall ->
if ("weather_of_location" == toolCall["function"]["name"].asText()) {
val toolArgs = objectMapper.readTree(toolCall["function"]["arguments"].asText())
val temperature = getTemperature(toolArgs["location"].asText())
val toolResponse = mapOf(
"tool_call_id" to toolCall["id"].asText(),
"role" to "tool",
"content" to temperature.toString()
)
allMessages.add(toolResponse)
}
}
}
val callInfo = CallInfo.from(
formattedPrompt.promptInfo,
startTime,
System.currentTimeMillis()
)
val sessionInfo = fpClient.sessions().create().sessionInfo
fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace)
.toolSchema(formattedPrompt.toolSchema)
)
} catch (e: Exception) {
throw RuntimeException("Failed to process JSON response", e)
}
}
.exceptionally { exception ->
System.err.println("Error: ${exception.message}")
exception.printStackTrace()
null
}
.join()
}
}
```
# SDKs
Source: https://docs.freeplay.ai/developer-resources/sdks
Native Freeplay SDKs for Python, TypeScript, and Java/Kotlin.
Freeplay provides native SDKs for the most popular development languages. Each SDK offers the same core functionality with language-idiomatic interfaces.
Using Claude Code or Cursor? Check out the experimental [Freeplay MCP server, skills, and plugin](/developer-resources/overview#claude-code--mcp-integration) to interact with Freeplay directly from your editor through natural language.
## Available SDKs
`pip install freeplay`
[PyPI](https://pypi.org/project/freeplay/) · [GitHub](https://github.com/freeplayai/freeplay-python)
`npm install freeplay`
[npm](https://www.npmjs.com/package/freeplay) · [GitHub](https://github.com/freeplayai/freeplay-node)
`ai.freeplay:client`
[Maven Central](https://central.sonatype.com/artifact/ai.freeplay/client)
## Core Capabilities
All SDKs provide:
* **Prompt Management**: Fetch and format prompts from Freeplay with variable interpolation
* **Observability**: Record sessions, traces, and completions automatically
* **Testing**: Create and execute test runs programmatically
* **Feedback**: Capture customer feedback on completions and traces
## Getting Started
For detailed setup instructions, configuration options, and usage examples, see the SDK documentation:
Complete guide to installing and configuring the Freeplay SDK
## SDK vs Framework Integrations
If you're using LangGraph, Vercel AI SDK, Google ADK, or other AI frameworks that emit clean OTel traces, consider our framework integrations or [Open Telemetry (OTel)](/developer-resources/integrations/tracing-with-otel) for automatic observability with minimal code changes.
| Feature | Freeplay SDK | Framework Integrations |
| ------------------ | ------------------- | ----------------------------- |
| Direct LLM control | Yes | Handled by framework |
| Automatic tracing | Manual | Automatic |
| Prompt management | Full support | Full support |
| Framework required | No | Yes (LangGraph, Vercel, etc.) |
| Best for | Custom integrations | Framework-specific projects |
# Calling Any Model
Source: https://docs.freeplay.ai/freeplay-sdk/calling-any-model
Record LLM interactions from any model or provider, including hosts or formats Freeplay doesn't natively support.
Freeplay allows you to record LLM interactions from any model or provider, including hosts or formats Freeplay doesn't natively support in our application ([see that list here](/account-setup/model-management#configuring-model-access)). When calling other models, you'll retrieve a Freeplay prompt template in our common format and reformat it as needed for the LLM you want to use.
Here is an example of calling Mistral 7B hosted on [BaseTen](https://www.baseten.co/)
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
from freeplay.llm_parameters import LLMParameters
import os
import requests
import time
# retrieve your prompt template from freeplay
prompt_vars = {"keyA": "valueA"}
prompt = fpClient.prompts.get(project_id=project_id,
template_name="template_name",
environment="latest")
# bind your variables to the prompt
formatted_prompt = prompt.bind(prompt_vars).format()
# customize the messages for your provider API
# In this case, mistral does not support system messages
# we need to merge the system message into the initial user message
messages = [{'role': 'user',
'content': formatted_prompt.messages[0]['content'] + '' + formatted_prompt.messages[1]['content']}]
# make your LLM call to your custom provider
# call mistral 7b hosted with baseten
s = time.time()
baseten_url = "https://model-xyz.api.baseten.co/production/predict"
headers = {
"Authorization": "Api-Key " + baseten_key,
}
data = {'messages': messages}
req = requests.post(
url=baseten_url,
headers=headers,
json=data
)
e = time.time()
# add the response to ongoing list of messages
resText = req.json()
messages.append({'role': 'assistant', 'content': resText})
# create a freeplay session
session = fpClient.sessions.create()
# Construct the CallInfo from scratch
call_info=CallInfo(
provider="mistral",
model="mistral-7b",
start_time=s,
end_time=e,
model_parameters=LLMParameters(
{"paramA": "valueA", "paramB": "valueB"}
),
usage=UsageTokens(prompt_tokens=123, completion_tokens=456)
)
# record the LLM interaction with Freeplay
payload = RecordPayload(
project_id=project_id,
all_messages=messages,
inputs=prompt_vars,
session_info=session.session_info,
prompt_version_info=prompt.prompt_info,
call_info=call_info
)
fpClient.recordings.create(payload)
```
```typescript typescript theme={null}
import Freeplay from "freeplay";
import axios from "axios";
// configure freeplay client
const fpClient = Freeplay({..});
/* PROMPT FETCH */
// set the prompt variables
let promptVars = {"keyA": "valueA"};
// fetch your prompt from freeplay
let promptTemplate = await fpClient.prompts.get({
projectId,
templateName: "template_name",
environment: "latest"
});
// format the prompt
let formattedPrompt = promptTemplate.bind(promptVars).format();
/* LLM CALL */
// build the messages for mistral
// mistral does not support system messages
// we need to merge the system message into the initial user message
let messages = [{'role': 'user', 'content': formattedPrompt.messages[0]['content'] + ' ' + formattedPrompt.messages[1]['content']}];
// configure baseten requests
const basetenUrl = "https://model-xyz.api.baseten.co/production/predict";
const headers = {
"Authorization": "Api-Key " + basetenKey
};
let data = {'messages': messages};
async function pingBaseTen(data) {
// send the requests
let start = new Date();
let responseData;
const response = await axios.post(basetenUrl, data, {headers: headers})
responseData = response.data;
console.log(responseData);
let end = new Date();
return [responseData, start, end];
}
let [responseData, start, end] = await pingBaseTen(data);
messages.push({'role': 'assistant', 'content': responseData});
console.log(messages);
/* RECORD */
// create a session
let session = fpClient.sessions.create({});
// log the interaction with freeplay
fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: promptVars,
sessionInfo: session,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: {provider: "mistral", model: "mistral-7B", startTime: start, endTime: end, modelParams: {'paramA': 'valueA', 'paramB': 'valueB'}}
});
```
```java java theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.recordings.*;
import ai.freeplay.client.thin.resources.recordings.CallInfo;
import ai.freeplay.client.thin.resources.sessions.SessionInfo;
import ai.freeplay.client.thin.resources.prompts.ChatMessage;
import ai.freeplay.example.java.ThinExampleUtils.Tuple2;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.List;
import java.net.http.HttpResponse;
import java.util.concurrent.CompletableFuture;
import static java.lang.String.format;
import static ai.freeplay.client.thin.Freeplay.Config;
import static java.net.http.HttpRequest.BodyPublishers.ofString;
public class CustomModel {
private static final ObjectMapper objectMapper = new ObjectMapper();
public static void main(String[] args) {
// set your environment variables
String basetenApiKey = System.getenv("BASETEN_KEY");
String freeplayApiKey = System.getenv("FREEPLAY_API_KEY");
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
String customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME");
// create your url from your customer domain
String baseUrl = format("https://%s.freeplay.ai/api", customerDomain);
// initialize the freeplay client
Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
);
// create the prompt variables
Map promptVars = Map.of("keyA", "varA");
/* FETCH PROMPT */
var promptFuture = fpClient.prompts().getFormatted(
projectId,
"template-name",
"latest",
promptVars
);
/* LLM CALL */
// set timer to be used to measure latency
long startTime = System.currentTimeMillis();
var llmFuture = promptFuture.thenCompose((FormattedPrompt formattedPrompt) ->
callMistral(
objectMapper,
basetenApiKey,
formattedPrompt.getPromptInfo().getModel(), // get the model name
formattedPrompt.getPromptInfo().getModelParameters(), // get the model parameters
formattedPrompt.getBoundMessages() // get the messages, formatted for your specified provider
).thenApply((HttpResponse response) ->
new Tuple2<>(formattedPrompt, response)
)
);
/* RECORD TO FREEPLAY */
var recordFuture = llmFuture.thenCompose((Tuple2, HttpResponse> promptAndResponse) ->
recordResult(
fpClient,
promptAndResponse.first, promptVars, startTime, promptAndResponse.second
)
);
recordFuture.thenApply(recordResponse -> {
System.out.println("Recorded call succeeded with completionId: " + recordResponse.getCompletionId());
return null;
}).exceptionally(exception -> {
System.out.println("Got exception: " + exception.getMessage());
return new RecordResponse(null);
}).join();
}
public static CompletableFuture recordResult(
Freeplay fpClient,
FormattedPrompt formattedPrompt,
Map promptVars,
long startTime,
HttpResponse response
) {
JsonNode bodyNode;
try{
bodyNode = objectMapper.readTree(response.body());
System.out.println(bodyNode);
} catch (JsonProcessingException e){
throw new RuntimeException("Unable to parse response body", e);
}
// add the returned message to the list of messages
String role = "assistant";
String content = bodyNode.asText();
List allMessages = formattedPrompt.allMessages(
new ChatMessage(role, content)
);
// construct the call info from scratch
CallInfo callInfo = new CallInfo(
"mistral",
"mistral-7b",
startTime,
System.currentTimeMillis(),
formattedPrompt.getPromptInfo().getModelParameters()
);
// construct the session info from the session object
SessionInfo sessionInfo = fpClient.sessions().create().getSessionInfo();
System.out.println("Completion: " + content);
// make the record call to freeplay
return fpClient.recordings().create(
new RecordInfo(projectId, allMessages)
.inputs(promptVars)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo)
);
}
public static CompletableFuture> callMistral(
ObjectMapper objectMapper,
String basetenApiKey,
String model,
Map llmParameters,
List messages
) {
try {
String basetenChatURL = "https://model-xyz.api.baseten.co/production/predict";
Map bodyMap = new LinkedHashMap<>();
bodyMap.put("messages", messages);
bodyMap.putAll(llmParameters);
String body = objectMapper.writeValueAsString(bodyMap);
HttpRequest.Builder requestBuilder = HttpRequest
.newBuilder(new URI(basetenChatURL))
.header("Content-Type", "application/json")
.header("Authorization", format("Api-Key %s", basetenApiKey))
.POST(ofString(body));
return HttpClient.newBuilder()
.build()
.sendAsync(requestBuilder.build(), HttpResponse.BodyHandlers.ofString());
} catch (Exception e) {
throw new RuntimeException(e);
}
}
}
```
```kotlin kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
import ai.freeplay.example.java.ThinExampleUtils.callOpenAI
import com.fasterxml.jackson.databind.ObjectMapper
import kotlinx.coroutines.future.await
import kotlinx.coroutines.runBlocking
// for baseten mistral call
import java.net.URI
import java.net.http.HttpClient
import java.net.http.HttpRequest
import java.net.http.HttpResponse
import java.net.http.HttpRequest.BodyPublishers
import java.util.concurrent.CompletableFuture
private val objectMapper = ObjectMapper()
fun callMistral(
objectMapper: ObjectMapper,
basetenApiKey: String,
model: String,
llmParameters: Map,
messages: List
): CompletableFuture> {
return try {
val basetenChatURL = "https://model-xyz.api.baseten.co/production/predict"
val bodyMap = LinkedHashMap()
bodyMap["messages"] = messages
bodyMap.putAll(llmParameters)
val body = objectMapper.writeValueAsString(bodyMap)
val requestBuilder = HttpRequest.newBuilder(URI(basetenChatURL))
.header("Content-Type", "application/json")
.header("Authorization", "Api-Key $basetenApiKey")
.POST(BodyPublishers.ofString(body))
HttpClient.newBuilder()
.build()
.sendAsync(requestBuilder.build(), HttpResponse.BodyHandlers.ofString())
} catch (e: Exception) {
throw RuntimeException(e)
}
}
fun main(): Unit = runBlocking {
// set your environment variables
val openaiApiKey = System.getenv("OPENAI_API_KEY")
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME")
val basetenApiKey = System.getenv("BASETEN_KEY")
// create your url from your customer domain
val baseUrl = String.format("https://%s.freeplay.ai/api", customerDomain)
// create the freeplay client
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
)
// set the prompt variables
val promptVars = mapOf("keyA" to "varA")
/* PROMPT FETCH */
val formattedPrompt = fpClient.prompts().getFormatted(
projectId,
"template-name",
"latest",
promptVars
).await()
/* LLM CALL */
// set timer to measure latency
val startTime = System.currentTimeMillis()
val llmResponse = callMistral(
objectMapper,
basetenApiKey,
formattedPrompt.promptInfo.model, // get the model name
formattedPrompt.promptInfo.modelParameters, // get the model params
formattedPrompt.getBoundMessages()
).await()
val endTime = System.currentTimeMillis()
// add the LLM response to the message set
val bodyNode = objectMapper.readTree(llmResponse.body())
println(bodyNode)
val role = "assistant"
val content = bodyNode.asText()
println("Completion: " + content)
val allMessage: List = formattedPrompt.allMessages(
ChatMessage(role, content)
)
/* RECORD */
// construct the call info from the prompt object
val callInfo = CallInfo(
"mistral", // hard code the provider
"mistral-7b", // hard code the model
startTime,
System.currentTimeMillis(),
formattedPrompt.promptInfo.modelParameters
)
// create the session and get the session info
val sessionInfo = fpClient.sessions().create().sessionInfo
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessage,
promptVars,
sessionInfo,
formattedPrompt.getPromptInfo(),
callInfo
)
).await()
println("Completion Record Succeeded with ${recordResponse.completionId}")
}
```
# Custom Model Parameters
Source: https://docs.freeplay.ai/freeplay-sdk/custom-model-parameters
Record additional model parameters beyond the defaults configured in the Freeplay UI.
The majority of critical model parameters like `temperature` and `max_tokens` can be configured in the Freeplay UI. However, if you are using additional parameters these can still be recorded during the Record call and will be displayed in the UI alongside the Completion.
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
# create a session which will create a UID
session = fpClient.sessions.create()
# build call info from scratch to log additional params
# get the base params
start_params = formatted_prompt.prompt_info.model_parameters
# set the additional parameters
additional_params = {"presence_penalty": 0.8, "n": 5}
# combine the two parameter sets
all_params = {**start_params, **additional_params}
call_info = CallInfo(
provider=formatted_prompt.prompt_info.provider,
model=formatted_prompt.prompt_info.model,
start_time=s,
end_time=e,
model_parameters=all_params # pass the full parameter set
)
# record the results
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session.session_info,
prompt_version_info=prompt_info,
call_info=call_info
)
completion_info = fpClient.recordings.create(payload)
```
```typescript typescript theme={null}
import Freeplay, { getSessionInfo } from "freeplay";
// create a session
let session = fpClient.sessions.create({});
// get the primary parameters
const baseParams = formattedPrompt.promptInfo.modelParameters;
// set the additional parameters
const additionalParams = { presence_penalty: 0.8, n: 5 };
// merge the parameters
const allParams = { ...baseParams, ...additionalParams };
console.log(allParams);
// construct the callInfo
const callInfoPayload = {
provider: formattedPrompt.promptInfo.provider,
model: formattedPrompt.promptInfo.model,
startTime: start,
endTime: end,
modelParameters: allParams,
};
// record the interaction with Freeplay
const completionInfo = await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: promptVars,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: callInfoPayload
});
```
```java java theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.recordings.*;
import ai.freeplay.client.thin.resources.sessions.SessionInfo;
import ai.freeplay.client.thin.resources.recordings.CallInfo;
// construct the session info from the session object
SessionInfo sessionInfo = fpClient.sessions().create().getSessionInfo();
// get the existing parameters
Map modelParams = formattedPrompt.getPromptInfo().getModelParameters();
// add the additional parameters
modelParams.put("presence_penalty", 0.8);
modelParams.put("n", 5);
// construct the call info
CallInfo callInfo = new CallInfo(
formattedPrompt.getPromptInfo().getProvider(),
formattedPrompt.getPromptInfo().getModel(),
startTime,
System.currentTimeMillis(),
modelParams
);
System.out.println("Completion: " + content);
// make the record call to freeplay
fpClient.recordings().create(
new RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo));
);
```
```kotlin kotlin theme={null}
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
/* RECORD */
// get the existing parameters
val modelParams = formattedPrompt.promptInfo.modelParameters;
modelParams["presence_penalty"] = 0.8;
modelParams["n"] = 5
// construct the call info from the prompt object
val callInfo = CallInfo(
formattedPrompt.promptInfo.provider,
formattedPrompt.promptInfo.model,
startTime,
endTime,
modelParams
)
// create the session and get the session info
val sessionInfo = fpClient.sessions().create().sessionInfo
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace)
).await()
```
# Customer Feedback & Events
Source: https://docs.freeplay.ai/freeplay-sdk/customer-feedback
Log customer feedback and events associated with LLM completions.
Freeplay lets you log customer feedback, client events, and any other customer experience-related metadata associated with any LLM completion. This can be useful to tie feedback from your application back to Freeplay creating a feedback loop, whether for explicit signals like a feedback score, or implicit signals like a request to regenerate a completion or edits to a draft.
**Customer Feedback supports arbitrary key-value pairs and accepts any string, boolean, integer or float**.
There is one special key-value pair to consider:
The key `freeplay_feedback` is a special case to capture your primary positive/negative user feedback signals from customers. It accepts only the following string values:
* `positive` will render a 👍 in the UI
* `negative` will render as a 👎 in the UI
## Methods Overview
| Method Name | Parameters | Description |
| ----------- | ---------------------------------------- | --------------------------------------------------- |
| `update` | `completion_id`: string `feedback`: dict | Log feedback associated with a specified completion |
## Log Customer Feedback
```python python theme={null}
# create a session which will create a UID
session = fpClient.sessions.create()
# record the results
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session.session_info,
prompt_version_info=prompt_info,
call_info=CallInfo.from_prompt_info(prompt_info, start_time=s, end_time=e),
)
# this will create the completion id needed for the logging of customer feedback
completion_info = fpClient.recordings.create(payload)
# add some customer feedback
fpClient.customer_feedback.update(
completion_id=completion_info.completion_id,
feedback={'freeplay_feedback': 'positive',
'link_clicked': True}
)
```
```typescript typescript theme={null}
// create a session
let session = fpClient.sessions.create({});
// record the interaction with Freeplay
const completionInfo = await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: promptVars,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
});
// record customer feedback
await fpClient.customerFeedback.update({
completionId: completionInfo.completionId,
customerFeedback: {
"freeplay_feedback": "positive",
"link_click": true
},
});
```
```java java theme={null}
fpClient.recordings().create(
new RecordInfo(projectId, allMessages)
.inputs(variables)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo))
.thenCompose(recordResponse ->
fpClient.customerFeedback().update(
projectId,
recordResponse.getCompletionId(),
Map.of("helpful", "thumbsup")
)
.thenApply(feedbackResponse -> recordResponse)
);
```
```kotlin kotlin theme={null}
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
).await()
// log customer feedback
val feedbackResponse = fpClient.customerFeedback().update(
recordResponse.completionId,
mapOf("freeplay_feedback" to "positive")
).await()
```
# Data Models
Source: https://docs.freeplay.ai/freeplay-sdk/data-models
Reference documentation for Freeplay SDK data models and payload structures.
Below are details of the Freeplay data model, which can be helpful to understand for more advanced usage. **Note the parameter names are written in snake case but some SDK languages make use of camel case instead**
## Record Payload
```python python theme={null}
from freeplay import RecordPayload
```
```typescript typescript theme={null}
import { RecordPayload } from "freeplay";
```
```java java theme={null}
import ai.freeplay.client.thin.resources.recordings.RecordPayload
```
| Parameter Name | Data Type | Description | Required |
| --------------------- | ---------------------- | -------------------------------------------------------------------------------- | -------- |
| project\_id | str | Freeplay's projectId | Y |
| all\_messages | List\[dict\[str, str]] | All messages in the conversation so far | Y |
| inputs | dict | The input variables | N |
| session\_info | SessionInfo | The session id for which the recording should be associated | N |
| prompt\_version\_info | PromptVersionInfo | The prompt info from a formatted prompt, used for version tracking | N |
| call\_info | CallInfo | Information associated with the LLM call | N |
| trace\_info | TraceInfo | The trace to associate this completion with (for agent workflows) | N |
| tool\_schema | List\[dict\[str, any]] | The tools/functions available to the model (for function calling/tool use) | N |
| test\_run\_info | TestRunInfo | Information associated with the Test Run if this recording is part of a Test Run | N |
## Record Response
The `recordings.create()` method returns a `RecordResponse` containing:
| Parameter Name | Data Type | Description |
| -------------- | --------- | ------------------------------------------------------------------------------------------------- |
| completion\_id | str | The unique ID of the recorded completion. Use this as `parent_id` when creating child tool spans. |
## Call Info
```python python theme={null}
from freeplay import CallInfo
```
```typescript typescript theme={null}
import { getCallInfo } from "freeplay";
```
```java java theme={null}
import ai.freeplay.client.thin.resources.recordings.CallInfo;
```
| Parameter Name | Data Type | Description | Required |
| ----------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| provider | string | The name of your LLM provider (e.g., "openai", "anthropic", "mistral") | N |
| model | string | The name of your model (e.g., "gpt-4o-mini", "claude-3-5-sonnet") | N |
| start\_time | float | The start time of the LLM call as Unix timestamp. This will be used to measure latency | N |
| end\_time | float | The end time of the LLM call as Unix timestamp. This will be used to measure latency | N |
| model\_parameters | LLMParameters | The parameters associated with your LLM call (e.g., temperature, max\_tokens) | N |
| usage | UsageTokens | Token count to record to Freeplay for cost calculation. If not included, Freeplay will estimate token counts using Tiktoken. | N |
| provider\_info | Dict\[str, Any] | Additional provider-specific informatio, e.g. azure\_deployment. | N |
| api\_style | "batch" or "default" | Set to "batch" if using a batch API (e.g., OpenAI Batch API) for accurate cost calculation. Use "default" or omit for standard API calls. | N |
## LLM Parameters
```python python theme={null}
from freeplay.llm_parameters import LLMParameters
```
```typescript typescript theme={null}
import { LLMParameters } from "freeplay";
```
```java java theme={null}
import ai.freeplay.client.thin.resources.recordings.LLMParameters
```
| Parameter Name | Data Type | Description | Required |
| -------------- | --------------- | ------------------------------------------------------------------- | -------- |
| members | Dict\[str, any] | Any parameters associated with your LLM call that you want recorded | Y |
## Trace Info
TraceInfo is returned by `session.create_trace()` and used to group completions together and track agent workflows. See [Traces](/freeplay-sdk/traces) for usage examples.
| Parameter Name | Data Type | Description | Required |
| ---------------- | ------------------------------ | ------------------------------------------------------------- | -------- |
| trace\_id | string | The unique ID of the trace | Auto |
| session\_id | string | The session this trace belongs to | Auto |
| input | str \| dict \| list | The input to the trace (user message or tool arguments) | N |
| agent\_name | string | Name of the agent for this trace | N |
| parent\_id | UUID | Parent trace or completion ID for creating nested hierarchies | N |
| kind | `'tool'` \| `'agent'` | Type of trace—use `'tool'` for tool execution spans | N |
| name | string | Name of the trace or tool | N |
| custom\_metadata | dict\[str, str \| int \| bool] | Custom metadata to associate with the trace | N |
## Test Run Info
```python python theme={null}
from freeplay import TestRunInfo
```
```typescript typescript theme={null}
import { getTestRunInfo } from "freeplay";
```
```java java theme={null}
import ai.freeplay.client.thin.resources.testruns.TestRun;
testRun.getTestRunInfo()
```
| Parameter Name | Data Type | Description | Required |
| -------------- | --------- | ------------------------ | -------- |
| test\_run\_id | string | The id of your Test Run | Y |
| test\_case\_id | string | The id of your Test Case | Y |
## OpenAI Function Call
```python python theme={null}
from freeplay.completions import OpenAIFunctionCall
```
```typescript typescript theme={null}
import { OpenAIFunctionCall } from "freeplay";
```
```java java theme={null}
import ai.freeplay.client.thin.resources.recordings.OpenAIFunctionCall
```
| Parameter Name | Data Type | Description | Required |
| -------------- | --------- | ------------------------------------------- | -------- |
| name | string | The name of the invoked function call | Y |
| arguments | string | The arguments for the invoked function call | Y |
# Organizing Principles
Source: https://docs.freeplay.ai/freeplay-sdk/organizing-principles
Understand the core namespaces and hierarchy that structure the Freeplay SDK.
The Freeplay SDK is organized around a clear hierarchy of entities that map to how generative AI applications work in practice. Understanding this structure will help you integrate Freeplay effectively.
## The Observability Hierarchy
Freeplay organizes your LLM application observability data in a three-level hierarchy:
```
Project
└── Session (a complete user interaction, multi-turn conversation, or instance of your application running)
└── Trace (optional grouping of related completions; required for complex agents)
└── Completion (a single LLM call)
```
* **Completions** are atomic LLM calls -- a prompt sent to a model and its response.
* **Traces** optionally group related completions, such as multiple LLM calls that power a single agent action. When building agents, Freeplay expects you to *name* traces to group related agent runs.
* **Sessions** contain all completions and traces for a logical user interaction (e.g., a chat conversation). You can choose to provide a session ID when recording completions. Sessions are created automatically if they don't exist.
For more detail on when to use traces vs. sessions, see [Sessions, Traces, and Completions](/core-concepts/observability/sessions-traces-and-completions).
## SDK Namespaces
All SDK operations are accessed through the Freeplay client object. The SDK is organized into namespaces that correspond to the core entities in Freeplay:
| Namespace | Purpose | Documentation |
| -------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------ |
| `client.sessions` | Create and manage sessions to group related completions | [Sessions](/freeplay-sdk/sessions) |
| `client.traces` | Create traces to group related completions within a session | [Traces](/freeplay-sdk/traces) |
| `client.recordings` | Record completions to Freeplay for observability | [Recording Completions](/freeplay-sdk/recording-completions) |
| `client.customer_feedback` | Log user feedback associated with completions | [Customer Feedback](/freeplay-sdk/customer-feedback) |
| `client.prompts` | Fetch and format prompt templates from Freeplay | [Prompts](/freeplay-sdk/prompts) |
| `client.test_runs` | Execute batch tests using saved datasets | [Test Runs](/freeplay-sdk/test-runs) |
Some SDKs use camelCase (TypeScript, Java, Kotlin) rather than snake\_case (Python) for method names, following language conventions. The functionality is identical across languages.
## Common Integration Flow
A basic integration follows this pattern:
1. **Fetch a specific version of a prompt template** from Freeplay with variables inserted by your code
2. **Call your LLM provider** (OpenAI, Anthropic, etc.) with the formatted prompt
3. **Record the completion** back to Freeplay for observability
```python python theme={null}
# 1. Fetch and format the prompt
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="my-prompt",
environment="prod",
variables={"question": user_question}
)
# 2. Call your LLM provider
response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
# 3. Record to Freeplay
session = fp_client.sessions.create()
fp_client.recordings.create(RecordPayload(
project_id=project_id,
all_messages=formatted_prompt.all_messages({
'role': response.choices[0].message.role,
'content': response.choices[0].message.content
}),
session_info=session.session_info,
prompt_version_info=formatted_prompt.prompt_info,
# ... additional parameters
))
```
For complete examples, see the [Recording Completions](/freeplay-sdk/recording-completions) page or browse [Code Recipes](/developer-resources/recipes/overview).
## Choosing Your Integration Approach
Freeplay offers multiple ways to integrate, depending on your needs:
| Approach | Language | Best For | Observability | Prompt Management |
| ------------------------------------------------------------------------ | ----------------------- | ------------------------------------ | ------------- | ----------------- |
| [**Freeplay SDK**](/freeplay-sdk/setup) | Python, TS, Java/Kotlin | Direct integration with full control | ✅ | ✅ |
| [**LangGraph**](/developer-resources/integrations/langgraph) | Python | LangGraph agent workflows | ✅ Auto | ✅ |
| [**Vercel AI SDK**](/developer-resources/integrations/vercel-ai-sdk) | TypeScript | TypeScript/JS AI applications | ✅ Auto | ✅ |
| [**Google ADK**](/developer-resources/integrations/adk) | Python | Google Agent Development Kit | ✅ Auto | ✅ |
| [**OpenTelemetry**](/developer-resources/integrations/tracing-with-otel) | Any | Any framework, standard tracing | ✅ | ❌ |
| [**HTTP API**](/api-reference) | Any | Custom implementations, automation | ✅ | ✅ |
OpenTelemetry integration provides observability only. For prompt management with OTel-traced applications, use the Freeplay SDK alongside your OTel instrumentation.
For framework integrations, see [AI Framework Integrations](/developer-resources/integrations/langgraph).
## Production Best Practices
Many Freeplay customers configure different client setups for different environments:
* **Development/Staging**: Fetch prompts from the Freeplay server for rapid iteration
* **Production**: Use [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling) to read prompts from local files for zero latency and resilience
See [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling) for implementation details.
## Next Steps
**Getting started:**
* [Setup](/freeplay-sdk/setup) - Install and configure the SDK
* [Prompts](/freeplay-sdk/prompts) - Fetch and format prompt templates
* [Recording Completions](/freeplay-sdk/recording-completions) - Log LLM interactions
**When you need more control:**
* [Sessions](/freeplay-sdk/sessions) - Explicitly create sessions with custom metadata (sessions are auto-created if you don't)
* [Traces](/freeplay-sdk/traces) - Group related completions for agent workflows with multiple LLM calls or tool invocations
# Prompts
Source: https://docs.freeplay.ai/freeplay-sdk/prompts
Retrieve and format prompt templates from Freeplay for use with any LLM provider.
Retrieve your prompt templates from the Freeplay server. All methods associated with your Freeplay prompt template are accessible from the `client.prompts` namespace.
## Methods Overview
Some SDKs will use camel case rather than snake case depending on convention for the given language
| Method Name | Parameters | Description |
| --------------- | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `get_formatted` | `project_id:` string `template_name:` string `environment:` string | Get a formatted prompt template object with variables inserted, messages formatted for the configured LLM provider (e.g. OpenAI, Anthropic), and model parameters for the LLM call. |
| `get` | `project_id:` string `template_name:` string `environment:` string | Get a prompt template by environment.*Note: The prompt template will not have variables substituted or be formatted for the configured LLM provider.* |
## Get a Formatted Prompt
Get a formatted prompt object with variables inserted, messages formatted for the associated LLM provider, and model parameters to use for the LLM call. This is the most convenient method for most prompt fetch use cases given that formatting is handled for you server side in Freeplay.
The `environment` parameter determines which version of your prompt is fetched. Default environment names are:
* `"prod"` - Production environment
* `"dev"` - Development environment
* `"latest"` - Latest version (use for development/testing)
Learn more about [managing prompts across environments](/core-concepts/prompt-management/managing-prompts) and [multi-environment configuration patterns](/core-concepts/prompt-management/prompt-bundling#multi-environment-configuration).
```python python theme={null}
# get a formatted prompt
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="template_name",
environment="latest",
variables={"keyA": "valueA"}
)
# Sample use in an LLM call
start = time.time()
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# add the response to your message set
all_messages = formatted_prompt.all_messages({
'role': chat_response.choices[0].message.role,
'content': chat_response.choices[0].message.content
})
```
```typescript typescript theme={null}
// set the prompt variables
let promptVars = {"keyA": "valueA"};
// fetch a formatted prompt template
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId: projectID,
templateName: "template_name",
environment: "latest",
variables: promptVars,
});
```
```java java theme={null}
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
// create the prompt variables
Map promptVars = Map.of("keyA", "valueA");
// get a formatted prompt
CompletableFuture> formattedPrompt = fpClient.prompts()
.getFormatted(projectId, "template_name", "latest", promptVars, null);
```
```kotlin kotlin theme={null}
// set the prompt variables
val promptVars = mapOf("keyA" to "valA")
/* PROMPT FETCH */
val formattedPrompt = fpClient.prompts().getFormatted(
projectId,
"template-name",
"latest",
promptVars
).await()
```
## Get a Prompt Template
Get a prompt template object that does not yet have variables inserted and has messages formatted in consistent LLM provider agnostic structure. It is particularly useful when you want to reuse a prompt template with different variables in the same execution path, like Test Runs.
This method gives you more control to handle formatting in your own code rather than server side in Freeplay, but requires a few more lines of code.
```python python theme={null}
# get an unformatted prompt template
template_prompt = fp_client.prompts.get(
project_id=project_id,
template_name="template_name",
environment="latest"
)
# to format the prompt
formatted_prompt = template_prompt.bind({"keyA": "valueA"}).format()
```
```typescript typescript theme={null}
// get a prompt template
let promptTemplate = await fpClient.prompts.get({
projectId: projectID,
templateName: "album_bot",
environment: "latest",
});
// format the prompt template
// set the prompt variables
let promptVars = { keyA: "valueA" };
// format the prompt template
let formattedPrompt = promptTemplate.bind(promptVars).format();
```
```java java theme={null}
import ai.freeplay.client.thin.resources.prompts.TemplatePrompt;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
// create the prompt variables
Map promptVars = Map.of("keyA", "valueA");
// get prompt template
CompletableFuture promptTemplate = fpClient.prompts().get(
projectId, "template_name", "environment");
// bind and format client side
FormattedPrompt
## Using History with Prompt Templates
`history` is a special object in Freeplay prompt templates for managing state over multiple LLM interactions by passing in previous messages. It accepts an array of prior messages when relevant.
Before using `history` in the SDK, you must configure it on your prompt template. See more details in the [Multi-Turn Chat Support section](/practical-guides/multi-turn-chat-support#history-and-prompt-templates).
Once you have `history` configured for a prompt template, you can pass it during the formatting process. The `history` messages will be inserted wherever you have your `history` placeholder in your prompt template.
```python python theme={null}
previous_messages = [{"role": "user", "content": "what are some dinner ideas..."},
{"role": "assistant", "content": "here are some dinner ideas..."}]
prompt_vars = {"question": "how do I make them healthier?"}
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="SamplePrompt",
environment="latest",
variables=prompt_vars,
history=previous_messages # pass the history messages here
)
# llm_prompt contains messages formatted for the provider
print(formatted_prompt.llm_prompt)
# output:
[
{'role': 'system', 'content': 'You are a polite assistant...'},
{'role': 'user', 'content': 'what are some dinner ideas...'},
{'role': 'assistant', 'content': 'here are some dinner ideas...'},
{'role': 'user', 'content': 'how do I make them healthier?'}
]
```
See a full implementation of using `history` in the context of a multi-turn chatbot application [here](/practical-guides/multi-turn-chat-support)
## Using Tool Schemas with Prompt Templates
You can define tool schemas alongside your prompt templates. The Freeplay SDK will format the tool schema based on the configured LLM provider, so you can pass the tool schema to the LLM provider as is.
```python python theme={null}
# get a formatted prompt
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="template_name",
environment="latest",
variables={"keyA": "valueA"}
)
# Sample use in an LLM call
start = time.time()
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
# Pass the tool schema to the LLM call
tool_schema=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
```
```typescript typescript theme={null}
// set the prompt variables
let promptVars = { keyA: "valueA" };
// fetch a formatted prompt template
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId: projectID,
templateName: "template_name",
environment: "latest",
variables: promptVars,
});
// Sample use in an LLM call
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
toolSchema: formattedPrompt.toolSchema,
...formattedPrompt.promptInfo.modelParameters,
});
```
```kotlin kotlin theme={null}
val variables = mapOf("keyA" to "valA")
fpClient.prompts()
.getFormatted>(
projectId,
"my-prompt",
"latest",
variables,
).thenCompose { formattedPrompt ->
val startTime = System.currentTimeMillis()
callAnthropic(
objectMapper,
anthropicApiKey,
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt,
formattedPrompt.systemContent.orElse(null),
formattedPrompt.toolSchema
).thenApply { Triple(formattedPrompt, it, startTime) }
}
.join()
```
## Record Evals From Your Code
Freeplay allows you to record client-side executed evals to Freeplay during the record step (more [here](/core-concepts/evaluations/code-evaluations)). Code evals are useful for running objective assertions or pairwise comparisons against ground truth data. They are passed as a key-value pair and can be associated with either a regular Session or a Test Run session.
```python python theme={null}
# record the results
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session.session_info,
prompt_version_info=prompt_info,
call_info=call_info,
eval_results={
"valid schema": True,
"valid category": True,
"string distance": 0.81
}
)
completion_info = fp_client.recordings.create(payload)
```
```typescript typescript theme={null}
// record the interaction with Freeplay
const completionInfo = await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: promptVars,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: callInfoPayload,
evalResults: {
"valid schema": true,
"valid category": true,
"string distance": 0.81,
},
});
```
```kotlin kotlin theme={null}
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
/* RECORD */
// get the existing parameters
val modelParams = formattedPrompt.promptInfo.modelParameters;
modelParams["presence_penalty"] = 0.8;
modelParams["n"] = 5
// construct the call info from the prompt object
val callInfo = CallInfo(
formattedPrompt.promptInfo.provider,
formattedPrompt.promptInfo.model,
startTime,
endTime,
modelParams
)
// create the session and get the session info
val sessionInfo = fpClient.sessions().create().sessionInfo
val evalResults = mapOf(
"valid schema" to true,
"valid category" to true,
"string distance" to 0.81
)
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessage,
).inputs(promptVars)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.promptInfo)
.callInfo(callInfo)
.evalResults(evalResults)
)
```
## Using your own IDs
Freeplay allows you to provide your own client-side UUIDs for both Sessions and Completions. This can be useful if you already have natural identifiers in your application code. Providing your own Completion Id also allows you to store the completion Id to be used for [recording customer feedback](/freeplay-sdk#customer-feedback--events) without having to wait for the record call to complete. Thus making the record call entirely non-blocking.
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo
from uuid import uuid4
## PROMPT FETCH
# set the prompt variables
prompt_vars = {"keyA": "valueA"}
# get a formatted prompt
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="template_name",
environment="latest",
variables=prompt_vars
)
## LLM CALL
# Make an LLM call to your provider of choice
start = time.time()
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# add the response to your message set
all_messages = formatted_prompt.all_messages({
'role': chat_response.choices[0].message.role,
'content': chat_response.choices[0].message.content
})
## RECORD
### CUSTOM IDS
# create your Ids (must be UUIDs)
session_id = uuid4()
completion_id = uuid4()
# Create sessionInfo with custom Ids
session_info = SessionInfo(
session_id=session_id,
custom_metadata={'keyA': 'valueA'}
)
### Extra data
call_info = CallInfo(
provider=formatted_prompt.prompt_info.provider,
model=formatted_prompt.prompt_info.model,
start_time=s,
end_time=e,
model_parameters=all_params # pass the full parameter set
)
# build the record payload
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session_info,
completion_id=completion_id,
prompt_version_info=formatted_prompt.prompt_info,
call_info=call_info,
)
# record the LLM interaction
fp_client.recordings.create(payload)
```
```typescript typescript theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
import { v4 as uuidv4 } from "uuid";
import OpenAI from "openai";
/* PROMPT FETCH */
// set the prompt variables
let promptVars = {"pop_star": "Taylor Swift"};
// fetch a formatted prompt template
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "template_name",
environment: "latest",
variables: promptVars,
});
/* LLM CALL */
// make the llm call
const openai = new OpenAI(process.env["OPENAI_API_KEY"]);
let start = new Date();
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters
});
let end = new Date();
console.log(chatCompletion.choices[0].message);
// update the messages
let messages = formattedPrompt.allMessages({
role: chatCompletion.choices[0].message.role,
content: chatCompletion.choices[0].message.content,
});
/* RECORD */
// create your own Ids
const sessionId = uuidv4();
const completionId = uuidv4();
// record the LLM interaction with Freeplay
await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: promptVars,
sessionInfo: {sessionId: sessionId, customMetadata: {"keyA": "valueA"}},
completionId: completionId,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end)
});
```
```kotlin kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
import ai.freeplay.example.java.ThinExampleUtils.callOpenAI
import com.fasterxml.jackson.databind.ObjectMapper
import java.util.UUID
import kotlinx.coroutines.future.await
import kotlinx.coroutines.runBlocking
private val objectMapper = ObjectMapper()
fun main(): Unit = runBlocking {
// set your environment variables
val openaiApiKey = System.getenv("OPENAI_API_KEY")
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME")
// create your url from your customer domain
val baseUrl = String.format("https://%s.freeplay.ai/api", customerDomain)
// create the freeplay client
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
)
// set the prompt variables
val promptVars = mapOf("keyA" to "valueA")
/* PROMPT FETCH */
val formattedPrompt = fpClient.prompts().getFormatted(
projectId,
"template-name",
"latest",
promptVars
).await()
/* LLM CALL */
// set timer to measure latency
val startTime = System.currentTimeMillis()
val llmResponse = callOpenAI(
objectMapper,
openaiApiKey,
formattedPrompt.promptInfo.model, // get the model name
formattedPrompt.promptInfo.modelParameters, // get the model params
formattedPrompt.getBoundMessages()
).await()
val endTime = System.currentTimeMillis()
// add the LLM response to the message set
val bodyNode = objectMapper.readTree(llmResponse.body())
val role = bodyNode.path("choices").path(0).path("message").path("role").asText()
val content = bodyNode.path("choices").path(0).path("message").path("content").asText()
println("Completion: " + content)
val allMessage: List = formattedPrompt.allMessages(
ChatMessage(role, content)
)
/* RECORD */
// construct the call info from the prompt object
val callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
endTime
)
// create the session and get the session info
val sessionInfo = fpClient.sessions().create().sessionInfo
// create your completion Id
val completionId = UUID.randomUUID()
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessage,
)
.completionId(completionId)
.inputs(promptVars)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.promptInfo)
.callInfo(callInfo)
.evalResults(evalResults)
)
).await()
println("Completion Record Succeeded with ${recordResponse.completionId}")
}
```
# Recording Completions
Source: https://docs.freeplay.ai/freeplay-sdk/recording-completions
Record LLM interactions to Freeplay for observability and evaluation.
Record LLM interactions to the Freeplay server for observability and evaluation. All methods associated with recordings are accessible via the `client.recordings` namespace.
## Methods Overview
| Method Name | Parameters | Returns | Description |
| ----------- | ----------------------------------------- | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| `create` | RecordPayload | RecordResponse with the new completion ID | Log your LLM interaction |
| `update` | `completion_id:` string, `feedback:` dict | — | Allows users to log additional feedback or information to a completion after it has already been recorded. |
## Record an LLM Interaction
Log your LLM interaction with Freeplay. This is assuming your have already retrieved a formatted prompt and made an LLM call as demonstrated in the [Prompts Section](/freeplay-sdk/prompts)
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo, CallInfo, UsageTokens
## PROMPT FETCH
# set the prompt variables
prompt_vars = {"keyA": "valueA"}
# get a formatted prompt
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="template_name",
environment="latest",
variables=prompt_vars
)
## LLM CALL
# Make an LLM call to your provider of choice
start = time.time()
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# add the response to your message set
all_messages = formatted_prompt.all_messages({
'role': chat_response.choices[0].message.role,
'content': chat_response.choices[0].message.content
})
## RECORD
# create a session
session = fp_client.sessions.create()
# build the record payload
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session.session_info,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted_prompt.prompt_info,
start_time=start,
end_time=end,
usage=UsageTokens(
prompt_tokens=chat_response.usage.prompt_tokens,
completion_tokens=chat_response.usage.completion_tokens
)
)
)
# record the LLM interaction
fp_client.recordings.create(payload)
```
```typescript typescript theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
import OpenAI from "openai";
/* PROMPT FETCH */
// set the prompt variables
let promptVars = {"pop_star": "Taylor Swift"};
// fetch a formatted prompt template
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "template_name",
environment: "latest",
variables: promptVars,
});
/* LLM CALL */
// make the llm call
const openai = new OpenAI(process.env["OPENAI_API_KEY"]);
let start = new Date();
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters
});
let end = new Date();
console.log(chatCompletion.choices[0].message);
// update the messages
let messages = formattedPrompt.allMessages({
role: chatCompletion.choices[0].message.role,
content: chatCompletion.choices[0].message.content,
});
/* RECORD */
// create a session
let session = fpClient.sessions.create({});
// record the LLM interaction with Freeplay
await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: promptVars,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
});
```
```java java theme={null}
package ai.freeplay.example.java;
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.recordings.*;
import ai.freeplay.client.thin.resources.sessions.SessionInfo;
import ai.freeplay.client.thin.resources.prompts.ChatMessage;
import ai.freeplay.example.java.ThinExampleUtils.Tuple2;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.util.Map;
import java.util.List;
import java.net.http.HttpResponse;
import java.util.concurrent.CompletableFuture;
import static java.lang.String.format;
import static ai.freeplay.client.thin.Freeplay.Config;
import static ai.freeplay.example.java.ThinExampleUtils.callOpenAI;
public class FreeplayRecord {
private static final ObjectMapper objectMapper = new ObjectMapper();
public static void main(String[] args) {
// set your environment variables
String openaiApiKey = System.getenv("OPENAI_API_KEY");
String freeplayApiKey = System.getenv("FREEPLAY_API_KEY");
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
String customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME");
// create your url from your customer domain
String baseUrl = format("https://%s.freeplay.ai/api", customerDomain);
// initialize the freeplay client
Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
);
// create the prompt variables
Map promptVars = Map.of("pop_star", "J Cole");
/* FETCH PROMPT */
var promptFuture = fpClient.prompts().getFormatted(
projectId,
"album_bot",
"latest",
promptVars
);
/* LLM CALL */
// set timer to be used to measure latency
long startTime = System.currentTimeMillis();
var llmFuture = promptFuture.thenCompose((FormattedPrompt formattedPrompt) ->
callOpenAI(
objectMapper,
openaiApiKey,
formattedPrompt.getPromptInfo().getModel(), // get the model name
formattedPrompt.getPromptInfo().getModelParameters(), // get the model parameters
formattedPrompt.getBoundMessages() // get the messages, formatted for your specified provider
).thenApply((HttpResponse response) ->
new Tuple2<>(formattedPrompt, response)
)
);
/* RECORD TO FREEPLAY */
var recordFuture = llmFuture.thenCompose((Tuple2, HttpResponse> promptAndResponse) ->
recordResult(
fpClient,
promptAndResponse.first, promptVars, startTime, promptAndResponse.second
)
);
recordFuture.thenApply(recordResponse -> {
System.out.println("Recorded call succeeded with completionId: " + recordResponse.getCompletionId());
return null;
}).exceptionally(exception -> {
System.out.println("Got exception: " + exception.getMessage());
return new RecordResponse(null);
}).join();
}
public static CompletableFuture recordResult(
Freeplay fpClient,
FormattedPrompt formattedPrompt,
Map promptVars,
long startTime,
HttpResponse response
) {
JsonNode bodyNode;
try{
bodyNode = objectMapper.readTree(response.body());
} catch (JsonProcessingException e){
throw new RuntimeException("Unable to parse response body", e);
}
// add the returned message to the list of messages
String role = bodyNode.path("choices").path(0).path("message").path("role").asText();
String content = bodyNode.path("choices").path(0).path("message").path("content").asText();
List allMessages = formattedPrompt.allMessages(
new ChatMessage(role, content)
);
// construct the call info from the formatted prompt
CallInfo callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
System.currentTimeMillis()
);
// construct the session info from the session object
SessionInfo sessionInfo = fpClient.sessions().create().getSessionInfo();
System.out.println("Completion: " + content);
// make the record call to freeplay
return fpClient.recordings().create(
new RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo)
);
}
}
}
```
```kotlin Kotlin theme={null}
package ai.freeplay.example.kotlin
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
import ai.freeplay.example.java.ThinExampleUtils.callOpenAI
import com.fasterxml.jackson.databind.ObjectMapper
import kotlinx.coroutines.future.await
import kotlinx.coroutines.runBlocking
private val objectMapper = ObjectMapper()
fun main(): Unit = runBlocking {
// set your environment variables
val openaiApiKey = System.getenv("OPENAI_API_KEY")
val freeplayApiKey = System.getenv("FREEPLAY_API_KEY")
val projectId = System.getenv("FREEPLAY_PROJECT_ID")
val customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME")
// create your url from your customer domain
val baseUrl = String.format("https://%s.freeplay.ai/api", customerDomain)
// create the freeplay client
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
)
// set the prompt variables
val promptVars = mapOf("keyA" to "valueA")
/* PROMPT FETCH */
val formattedPrompt = fpClient.prompts().getFormatted(
projectId,
"template-name",
"latest",
promptVars
).await()
/* LLM CALL */
// set timer to measure latency
val startTime = System.currentTimeMillis()
val llmResponse = callOpenAI(
objectMapper,
openaiApiKey,
formattedPrompt.promptInfo.model, // get the model name
formattedPrompt.promptInfo.modelParameters, // get the model params
formattedPrompt.getBoundMessages()
).await()
val endTime = System.currentTimeMillis()
// add the LLM response to the message set
val bodyNode = objectMapper.readTree(llmResponse.body())
val role = bodyNode.path("choices").path(0).path("message").path("role").asText()
val content = bodyNode.path("choices").path(0).path("message").path("content").asText()
println("Completion: " + content)
val allMessage: List = formattedPrompt.allMessages(
ChatMessage(role, content)
)
/* RECORD */
// construct the call info from the prompt object
val callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
endTime
)
// create the session and get the session info
val sessionInfo = fpClient.sessions().create().sessionInfo
// build the final record payload
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace)
).await()
println("Completion Record Succeeded with ${recordResponse.completionId}")
}
```
### Using OpenAI or Anthropic's Batch APIs
Some LLM providers offer a batch method of generating completions. If you are using the Batch API, you can log your results to Freeplay with `api_style="batch"`. This parameter is needed to calculate accurate costs for batch API usage, which are often significantly lower than regular completions. The general flow for logging batch data looks like:
1. Create batch file with Freeplay completion tracking For each input, format your prompt using Freeplay's template, then create a completion record with `api_style='batch'` in the CallInfo. Store the returned `completion_id` to update the completion in Freeplay once the batch request has completed.
2. Submit batch to OpenAI and poll for completion Upload your batch file to OpenAI, create the batch request, and poll until the batch status is "completed".
3. Update Freeplay with results Once complete, read the batch output file and update each Freeplay completion using the `completion_id` from the response to match it back to the original request.
[Here](/developer-resources/recipes/openai-batch-api) is a full code example for reference.
*Note: Only use the batch api\_style if you are calling a batch LLM API. Setting this parameter incorrectly may result in incorrect costs displaying in Freeplay!*
```python python theme={null}
CallInfo(provider="openai", model="gpt-4o-mini",
usage=UsageTokens(prompt_tokens=123, completion_tokens=456), api_style='batch'
)
```
## Recording Tools
You can record tool calls and their associated schemas for both OpenAI and Anthropic. These recorded completions and tool schemas can be viewed in the observability tab.
When using function calling or tool use, pass the `tool_schema` parameter in your `RecordPayload`.
This should be a list of tool/function definitions that were available to the model. The schema can be retrieved from
`formatted_prompt.tool_schema` if defined in your prompt template, or you can pass your own tool definitions in the same
format you use to pass them to the model.
The example below shows the **default approach**: tool calls are recorded as part of the completion messages. When you call `formatted_prompt.all_messages()`, the LLM's tool call output and subsequent tool results are concatenated into the message history alongside other messages.
For more granular tool observability, you can also create **explicit tool spans** using traces with `kind='tool'`. This renders tool arguments and results as separate spans in the trace view, which is useful for debugging complex agent workflows. See [Tool Calls](/practical-guides/tools#logging-tool-calls-to-freeplay) for both approaches.
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
from openai import OpenAI
## FETCH PROMPT
# get your formatted prompt from freeplay including any associated tool schemas
question = "What is the latest AI news?"
prompt_vars = {"question": question}
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="NewsSummarizerFuncEnabled",
variables=prompt_vars,
environment="latest"
)
## LLM CALL
# make your llm call with your tool schemas passed in
s = time.time()
openai_client = OpenAI(api_key=openai_key)
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters,
tools=formatted_prompt.tool_schema
)
e = time.time()
## RECORD
# create a session
session = fp_client.sessions.create()
# Append the response to the messages
messages = formatted_prompt.all_messages(chat_response.choices[0].message)
# Optionally prep Freeplay parameters
# Provide token use
token_usage = UsageTokens(
prompt_tokens=chat_response.usage.prompt_tokens,
completion_tokens=chat_response.usage.completion_tokens
)
# Provide timing information
call_info = CallInfo.from_prompt_info(
prompt_info=formatted_prompt.prompt_info,
start_time=start,
end_time=end,
usage=usage
)
## RECORD
# Record to Freeplay
record_response = fpclient.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=prompt_vars,
prompt_version_info=formatted_prompt.prompt_info,
call_info=call_info,
tool_schema=formatted_prompt.tool_schema # Optionally record the tool schema as well
)
)
```
```typescript typescript theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
/* FETCH PROMPT */
// set the prompt variables
let promptVars = {"question": "What is the latest AI news?"};
// fetch a formatted prompt template
let formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "NewsSummarizerFuncEnabled",
environment: "latest",
variables: promptVars,
});
/* LLM CALL */
const openai = new OpenAI(process.env["OPENAI_API_KEY"]);
let start = new Date();
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
tools: formattedPrompt.tool_schema,
...formattedPrompt.promptInfo.modelParameters
});
let end = new Date();
console.log(chatCompletion.choices[0]);
/* RECORD */
// create the session
let session = fpClient.sessions.create({});
// Append the response to the messages
let messages = formattedPrompt.allMessages(chatCompletion.choices[0].message);
fpClient.recordings.create({
projectId,
// Last message captures the tool call if it exists
allMessages: messages,
sessionInfo: session,
inputs: promptVars,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: CallInfo.fromPromptInfo(formattedPrompt.promptInfo, start, end),
// Records the associated tool schema
toolSchema: formattedPrompt.toolSchema
});
```
```kotlin kotlin theme={null}
import ai.freeplay.client.thin.Freeplay
import ai.freeplay.client.thin.resources.prompts.ChatMessage
import ai.freeplay.client.thin.resources.recordings.CallInfo
import ai.freeplay.client.thin.resources.recordings.RecordInfo
object OpenAIToolsExample {
@JvmStatic
fun main(args: Array) {
/* FETCH PROMPT */
// set the prompt variables
val variables = mapOf("question" to "What is the latest AI news?")
fpClient.prompts()
.getFormatted>(
projectId,
"NewsSummarizerFuncEnabled",
"latest",
variables,
null
).thenCompose { formattedPrompt ->
/* LLM CALL */
val startTime = System.currentTimeMillis()
callOpenAIWithTools(
openaiApiKey,
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt,
formattedPrompt.toolSchema
).thenApply { response ->
Triple(formattedPrompt, response, startTime)
}
}.thenCompose { (formattedPrompt, bodyNody, startTime) ->
// messageNode is nested Object which is Map
val messageNode = bodyNode["choices"][0]["message"]
// Append the response to the messages
val allMessages = formattedPrompt.allMessages(message).toMutableList()
val callInfo = CallInfo.from(
formattedPrompt.promptInfo,
startTime,
System.currentTimeMillis()
)
/* RECORD */
// create the session
val sessionInfo = fpClient.sessions().create().sessionInfo
fpClient.recordings().create(
RecordInfo(
projectId,
allMessages,
).inputs(variables)
.sessionInfo(session.sessionInfo)
.promptVersionInfo(prompt.promptInfo)
.callInfo(callInfo)
.traceInfo(trace).toolSchema(formattedPrompt.toolSchema)
).await()
}
.join()
}
}
```
### Updating a Completion
Freeplay allows you to update a completion once it has already been recorded. This can be useful to add client evals or additional messages to the completion. To do this, you will need to have the `project_id` and `completion_id`. The completion\_id is returned via `recordings.create`. The code example below shows this in action:
```python python theme={null}
#############################################################################
# UPDATE THE COMPLETION - Use the record_response to get the completion_id
#############################################################################
from freeplay.resources.recordings import RecordUpdatePayload
# Update the completion with customer feedback
# This uses the update method to attach feedback to a specific completion
fp_client.recordings.update(
RecordUpdatePayload(
project_id=PROJECT_ID,
completion_id=final_completion_id,
eval_results={
"accuracy": accuracy,
"precision": precision,
"recall": recall,
"f1": f1,
},
)
)
```
## Calling Any Model
Freeplay allows you to record LLM interactions from any model or provider, including hosts or formats Freeplay doesn't natively support. See [Calling Any Model](/freeplay-sdk/calling-any-model) for full examples across all supported languages.
## Custom Model Parameters
You can record additional model parameters beyond the defaults configured in the Freeplay UI. See [Custom Model Parameters](/freeplay-sdk/custom-model-parameters) for full examples.
# Sessions
Source: https://docs.freeplay.ai/freeplay-sdk/sessions
Create and manage sessions to group related LLM completions together.
Sessions are the top-level grouping in Freeplay's observability hierarchy. Every completion you record belongs to a session—you provide a session ID when recording, and the session is created automatically if it doesn't already exist.
## Understanding Sessions
A session represents a logical user interaction or instance of your application running. Think of it as a container that groups related completions together for analysis.
The `sessions.create()` method generates a session ID (UUID v4) and returns a Session object. You can also optionally set custom metadata at creation time. While you could generate your own UUID and let the session be created implicitly when recording your first completion, using `sessions.create()` is the recommended pattern.
**Create a new session when:**
* A user starts a new conversation
* A distinct workflow or task begins
* A new instance of your application starts handling requests
For more context on sessions vs. traces, see [Sessions, Traces, and Completions](/core-concepts/observability/sessions-traces-and-completions).
## Methods Overview
| Method | Description | Python | TypeScript | Java/Kotlin |
| ------ | -------------------- | ----------------------------------------- | --------------------------------------- | ----------------------------------------- |
| Create | Create a new session | `sessions.create()` | `sessions.create({})` | `sessions().create()` |
| Delete | Delete a session | `sessions.delete(project_id, session_id)` | `sessions.delete(projectId, sessionId)` | `sessions().delete(projectId, sessionId)` |
## Create a Session
```python python theme={null}
# create a session
session = fp_client.sessions.create()
```
```typescript typescript theme={null}
// create a session
const session = fpClient.sessions.create({});
```
```java java theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.sessions.Session;
// create a session
Session session = fpClient.sessions().create();
```
```kotlin kotlin theme={null}
val session = fpClient.sessions().create()
```
## Custom Metadata
Freeplay enables you to log Custom Metadata associated with any Session. This is fully customizable and can take any arbitrary key-value pairs.
The `custom_metadata` parameter accepts a dictionary (Python) or object (TypeScript/JavaScript) with arbitrary key-value pairs. Use Custom Metadata to record things like customer ID, thread ID, RAG version, Git commit SHA, or any other information you want to track.
```python python theme={null}
# create a session with custom metadata
session = fp_client.sessions.create(
custom_metadata={
"keyA": "valueA",
"keyB": False
}
)
```
```typescript typescript theme={null}
// create a session with custom metadata
const session = fpClient.sessions.create({
keyA: "valueA",
keyB: "valueB"
});
```
```java java theme={null}
import ai.freeplay.client.thin.resources.sessions.Session;
// create a session with Custom Metadata
Session sessionB = fpClient.sessions().create().customMetadata(Map.of("keyA", "valueA"));
```
```kotlin kotlin theme={null}
val session = fpClient.sessions().create()
.customMetadata(mapOf("keyA" to "valueA"))
```
## Delete a Session
```python python theme={null}
project_id = 'bf56b063-80dc-4ad5-91f6-f7067ad1fa06'
session_id = session.session_id
fp_client.sessions.delete(project_id, session_id)
```
```typescript typescript theme={null}
const projectId = 'bf56b063-80dc-4ad5-91f6-f7067ad1fa06'
const sessionId = session.sessionId;
await fpClient.sessions.delete(projectId, sessionId);
```
```java java theme={null}
String projectId = "bf56b063-80dc-4ad5-91f6-f7067ad1fa06";
SessionInfo sessionInfo = fpClient.sessions().create().getSessionInfo();
fpClient.sessions().delete(projectId, sessionInfo.getSessionId());
```
## Using Sessions with Completions
When you record a completion, pass the `session_info` from your session to associate the completion with that session:
```python python theme={null}
session = fp_client.sessions.create()
# ... make LLM call ...
fp_client.recordings.create(RecordPayload(
project_id=project_id,
session_info=session.session_info,
# ... other parameters
))
```
```typescript typescript theme={null}
const session = fpClient.sessions.create({});
// ... make LLM call ...
await fpClient.recordings.create({
projectId,
sessionInfo: session.sessionInfo,
// ... other parameters
});
```
For complete recording examples, see [Recording Completions](/freeplay-sdk/recording-completions).
# Setup
Source: https://docs.freeplay.ai/freeplay-sdk/setup
Install and configure the Freeplay SDK for Python, Node.js, or Java/Kotlin.
**Using a coding agent?** Point it to [docs.freeplay.ai/llms.txt](https://docs.freeplay.ai/llms.txt) for LLM-optimized documentation.
# Installation
We offer the SDK in several popular programming languages.
```python python theme={null}
pip install freeplay
```
```node node theme={null}
npm install freeplay
```
```xml java theme={null}
ai.freeplayclientx.x.xxcom.google.cloudgoogle-cloud-vertexai1.5.0
```
```kotlin kotlin theme={null}
// Add the Freeplay SDK to your build.gradle.kts
dependencies {
implementation("ai.freeplay:client:x.x.xx")
}
```
The Python and TypeScript SDKs are open source. View the source, report issues, or contribute on GitHub:
[freeplay-python](https://github.com/freeplayai/freeplay-python) ·
[freeplay-node](https://github.com/freeplayai/freeplay-node)
# Freeplay Client
The Freeplay client object will be your point of entry for all SDK usage. For most users, the base freeplay domain will be `https://app.freeplay.ai` with api exposed at `https://app.freeplay.ai/api`.
## Custom Domain
If your organization has been assigned a custom Freeplay domain it will look like this: `https://acme.freeplay.ai` and your SDK connection URL will have an additional `/api` appended to it (e.g., `https://acme.freeplay.ai/api` or `https://app.freeplay.ai/api`).
The `api_base` (Python) or `baseUrl` (Node.js) parameter is **required** and must match your Freeplay instance URL. Make sure to include the `/api` suffix for all domains.
Note the rest of the documentation will use `https://app.freeplay.ai`.
## Authentication
Freeplay authenticates your API request using your API Key which can be managed through the Freeplay application at `https://app.freeplay.ai/settings/api-access`
## Client Instantiation
The first step to using Freeplay is to create a client
```python python theme={null}
from freeplay import Freeplay
import os
# create a freeplay client object
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api"
)
```
```node node theme={null}
import Freeplay from "freeplay";
// create your freeplay client
const fpClient = new Freeplay({
freeplayApiKey: process.env["FREEPLAY_API_KEY"],
baseUrl: "https://app.freeplay.ai/api",
});
```
```java java theme={null}
import ai.freeplay.client.thin.Freeplay;
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
String customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME");
Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
);
```
```kotlin kotlin theme={null}
// create the freeplay client
val fpClient = Freeplay(
Freeplay.Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
)
```
**Production Setup**: For production deployments, consider using [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling) to fetch prompts from local files instead of the server. Many teams configure separate clients for different environments—server-based fetching for dev/staging, and bundled prompts for production.
For API rate limits and payload constraints, see [Platform Limits](/openapi/limits).
# Test Runs
Source: https://docs.freeplay.ai/freeplay-sdk/test-runs
Run batch tests of your LLM prompts and chains programmatically.
Test Runs in Freeplay provide a structured way for you to run batch tests of your LLM prompts and chains. All methods associated with the Test Runs concept in Freeplay are accessible via the `client.test_runs` namespace. Test runs can be completed using Completion or Trace datasets. We will focus on code in this section, but for more detail on the Test Runs concept see [Test Runs](/core-concepts/test-runs/test-runs).
## Methods Overview
| Method Name | Parameters | Description |
| ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `create` | `project_id:` string `testlist:` string (dataset name) `include_outputs:` bool (optional, defaults to False) `name:` string (optional) `description:` string (optional) | Instantiate a Test Run object server side and get an Id to reference your Test Run instance. To get expected outputs with your test cases, set `include_outputs=True`. |
The `testlist` parameter accepts the name of a dataset stored in Freeplay. This parameter name is preserved for backwards compatibility. In the UI and documentation, we use "dataset" to refer to this concept.
## Step by Step Usage
### Create a new Test Run
```python python theme={null}
from freeplay import Freeplay, RecordPayload, TestRunInfo
from openai import OpenAI
# create a new test run
test_run = fpClient.test_runs.create(
project_id=project_id,
testlist= # TODO fill in with the name of the dataset stored in Freeplay
name="mytestrun", # Name of the test run in Freeplay
description="this is a test test!"
)
```
```typescript typescript theme={null}
import Freeplay, { getSessionInfo, getCallInfo, getTestRunInfo} from "freeplay";
// create a test run
const testRun = await fpClient.testRuns.create({
projectId: fpProjectId,
testList: 'test-list-name',
name: 'my test run',
description: 'this is a test test'
});
```
```java java theme={null}
// see full implmentation in Iterate over each Test Case
```
```kotlin kotlin theme={null}
val testRun = fpClient.testRuns().create(
projectId,
"test-list",
"my test run", "this is a test test!"
).await()
```
### Retrieve your Prompts
Retrieve the prompts needed for your Test Run
```python python theme={null}
# get the prompt associated with the test run
template_prompt = fpClient.prompts.get(
project_id=project_id,
template_name="template-name",
environment="latest"
)
```
```typescript typescript theme={null}
// fetch the prompt template for the test run
let templatePrompt = await fpClient.prompts.get({
projectId: fpProjectId,
templateName: "template-name",
environment: "latest",
});
```
```java java theme={null}
// see full implmentation in Iterate over each test Case
```
```kotlin kotlin theme={null}
val templatePrompt = fpClient.prompts().get(
projectId,
"template-name",
"prod").await()
```
### Iterate over each Test Case
For the code you want to test: loop over each test case from the dataset, make an LLM call, and record the results with a link to your test run.
```python python theme={null}
# iterate over each test case
for test_case in test_run.test_cases:
# format the prompt with the test case variables
formatted_prompt = template_prompt.bind(test_case.variables).format()
# make your llm call
s = time.time()
openai_client = OpenAI(api_key=openai_key)
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
e = time.time()
# append the results to the messages
all_messages = formatted_prompt.all_messages({
'role': chat_response.choices[0].message.role,
'content': chat_response.choices[0].message.content
})
call_info = CallInfo.from_prompt_info(
formatted_prompt.prompt_info,
start_time=s,
end_time=e,
usage=UsageTokens(
chat_response.usage.prompt_tokens,
chat_response.usage.completion_tokens
)
)
# create a session which will create a UID
session = fp_client.sessions.create()
# build the record payload
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=test_case.variables, # Variables from the test case are the inputs
session_info=session.session_info,
# IMPORTANT: link the record call to the test run and test case
test_run_info=test_run.get_test_run_info(test_case.id),
prompt_version_info=formatted_prompt.prompt_info, # log the prompt information
call_info=call_info
)
# record the results to freeplay
fpClient.recordings.create(payload)
```
```typescript typescript theme={null}
for (const testCase of testRun.testCases) {
// create a formatted prompt from the test case
const formattedPrompt = templatePrompt.bind(testCase.variables).format();
// make the llm call
let start = new Date();
const chatCompletion = await openai.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
...formattedPrompt.promptInfo.modelParameters
});
let end = new Date();
console.log(chatCompletion.choices[0].message);
// update the messages
let messages = formattedPrompt.allMessages({
role: chatCompletion.choices[0].message.role,
content: chatCompletion.choices[0].message.content,
});
// create a session
let session = fpClient.sessions.create({});
// record the test case interaction with Freeplay
await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: testCase.variables,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
testRunInfo: getTestRunInfo(testRun, testCase.id)
});
}
```
```java java theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.ChatMessage;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import ai.freeplay.client.thin.resources.prompts.TemplatePrompt;
import ai.freeplay.client.thin.resources.recordings.CallInfo;
import ai.freeplay.client.thin.resources.recordings.RecordInfo;
import ai.freeplay.client.thin.resources.recordings.RecordResponse;
import ai.freeplay.client.thin.resources.sessions.SessionInfo;
import ai.freeplay.client.thin.resources.testruns.TestCase;
import ai.freeplay.client.thin.resources.testruns.TestRun;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.net.http.HttpResponse;
import java.util.List;
import java.util.concurrent.CompletableFuture;
import java.util.concurrent.ExecutionException;
import static ai.freeplay.example.java.ThinExampleUtils.callOpenAI;
import static java.lang.String.format;
import static ai.freeplay.client.thin.Freeplay.Config;
import static java.util.stream.Collectors.toList;
public class TestClass {
private static final ObjectMapper objectMapper = new ObjectMapper();
public static void main(String[] args) throws ExecutionException, InterruptedException {
// set your environment variables
String openaiApiKey = System.getenv("OPENAI_API_KEY");
String freeplayApiKey = System.getenv("FREEPLAY_API_KEY");
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
String customerDomain = System.getenv("FREEPLAY_CUSTOMER_NAME");
// create your url from your customer domain
String baseUrl = format("https://%s.freeplay.ai/api", customerDomain);
Freeplay fpClient = new Freeplay(Config()
.freeplayAPIKey(freeplayApiKey)
.customerDomain(customerDomain)
);
// create the test run
List recordResponses = fpClient.testRuns().create(projectId, "test-list-name")
.thenCompose(testRun ->
fpClient.prompts().get(projectId, "template-name", "latest")
.thenCompose(templatePrompt -> {
var futures =
testRun.getTestCases().stream()
.map(testCase ->
handleTestCase(
fpClient,
openaiApiKey,
templatePrompt,
testRun,
testCase))
.collect(toList());
return CompletableFuture.allOf(futures.toArray(CompletableFuture[]::new))
.thenApply((Void v) -> futures.stream().map(CompletableFuture::join).collect(toList()));
})
).join();
System.out.println("Record responses: " + recordResponses);
}
private static CompletableFuture handleTestCase(
Freeplay fpClient,
String openaiApiKey,
TemplatePrompt templatePrompt,
TestRun testRun,
TestCase testCase
) {
FormattedPrompt formattedPrompt =
templatePrompt.bind(testCase.getVariables()).format();
long startTime = System.currentTimeMillis();
return callOpenAI(
objectMapper,
openaiApiKey,
formattedPrompt.getPromptInfo().getModel(),
formattedPrompt.getPromptInfo().getModelParameters(),
formattedPrompt.getBoundMessages()
).thenCompose((HttpResponse response) ->
recordOpenAI(fpClient, testRun, testCase, formattedPrompt, startTime, response)
);
}
private static CompletableFuture recordOpenAI(
Freeplay fpClient,
TestRun testRun,
TestCase testCase,
FormattedPrompt formattedPrompt,
long startTime,
HttpResponse response
) {
JsonNode bodyNode;
try {
bodyNode = objectMapper.readTree(response.body());
} catch (JsonProcessingException e) {
throw new RuntimeException("Unable to parse response body.", e);
}
// add the returned message to the list of messages
String role = bodyNode.path("choices").path(0).path("message").path("role").asText();
String content = bodyNode.path("choices").path(0).path("message").path("content").asText();
List allMessages = formattedPrompt.allMessages(
new ChatMessage(role, content)
);
CallInfo callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
System.currentTimeMillis()
);
SessionInfo sessionInfo = fpClient.sessions().create().getSessionInfo();
System.out.println("Completion: " + content);
return fpClient.recordings().create(
new RecordInfo(projectId, allMessages)
.inputs(testCase.getVariables())
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo)
.toolSchema(formattedPrompt.getToolSchema())
.testRunInfo(testRun.getTestRunInfo(testCase.getTestCaseId())));
}
}
```
```kotlin kotlin theme={null}
for (testCase in testRun.testCases) {
// format the prompt with test case variables
val formattedPrompt = templatePrompt.bind(testCase.variables).format()
// make your llm call
val startTime = System.currentTimeMillis()
val llmResponse = callOpenAI(
objectMapper,
anthropicApiKey,
formattedPrompt.promptInfo.model,
formattedPrompt.promptInfo.modelParameters,
formattedPrompt.formattedPrompt
).await()
val bodyNode = objectMapper.readTree(llmResponse.body())
println("Recording the result")
// append the results to your message set
val allMessages = formattedPrompt.allMessages(
ChatMessage("Assistant", bodyNode.path("completion").asText())
)
val callInfo = CallInfo.from(
formattedPrompt.getPromptInfo(),
startTime,
System.currentTimeMillis()
)
// create a session
val sessionInfo = fpClient.sessions().create().sessionInfo
// record the test case results
val recordResponse = fpClient.recordings().create(
RecordInfo(
projectId,
allMessages
).inputs(testCase.variables)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo)
.testRunInfo(testRun.getTestRunInfo(testCase.testCaseId))
).await()
println("Recorded with completionId ${recordResponse.completionId}")
}
```
## Agent Test Runs
To execute tests in code, iterate through each test case in your dataset and run it through your full agent workflow. This end-to-end testing approach is particularly valuable for agentic systems, where the goal is to observe how changes—whether to prompts, tools, or orchestration logic—affect the final output.
Trace-level tests allow you to simulate production-like behavior and evaluate the agent holistically. As each test case runs, its input is passed into your system, and the resulting trace is logged for evaluation and analysis. See the full example [here](/core-concepts/test-runs/end-to-end-test-runs).
Below is an **psuedo code** example to show the general logic for test-runs:
```python python theme={null}
##################################
# Preare the test
##################################
trace_dataset = "Your dataset name that targets an agent"
test_name = "Name of the test you are running"
# Create the test
test_run = fp_client.test_runs.create(
project_id=project_id,
testlist=trace_dataset,
name=test_name
)
##################################
# Initialize your code or agents
##################################
"""
NOTE: It is important that all agents get passed the test_run_info and it is used when RecordPayload is called. This is how Freeplay tracks the test run
"""
my_agent = MyAgent()
################################################
# Loop over all test cases and execute system code
################################################
for test_case in test_run.trace_test_cases:
# Initialize a Freeplay Session
session = fp_client.sessions.create()
# NOTE: Get the test case information; This must be passed with any record calls
test_case_info = test_run.get_test_run_info(test_case.id)
# Create your Agent
question = test_case.input
trace_info = session.create_trace(
input=question,
agent_name="MyAgent", # Agent's name
custom_metadata={
"version": "1.0.0" # custom dict[str, str]
}
)
################################################
# Execute Agent
################################################
# Run your system code using the agents inputs/outputs
# This is where all the sub-agent processes and recordings happen
MyAgent.run(
input=question,
test_run_info=test_run_info # NOTE: This must be passed to record calls
)
# Record the agent output, pass the test run info here as well
trace_info.record_output(
project_id,
completion.choices[0].message.content,
eval_results={
'evaluation_score': 0.48,
'is_high_quality': True
},
test_run_info=test_run_info
)
```
# Traces
Source: https://docs.freeplay.ai/freeplay-sdk/traces
Group completions within sessions using traces to track agent workflows.
## Using Traces
Traces are an organizing component of a session. They are used to group completions together and provide a way to track the progress of a session.
Named traces form the basis of Freeplay's support for Agents. For a comprehensive guide on building agents with traces, see [Agents](/practical-guides/agents).
| Method Name | Parameters | Description |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| `session.create_trace ` | `input`: str - Input to the trace. `agent_name`: str (optional) - Name of the agent/trace. `parent_id`: UUID (optional) - Parent trace or completion ID for nesting. `kind`: str (optional) - Either `'tool'` or `'agent'`. `name`: str (optional) - Name of the trace/tool. `custom_metadata`: dict\[str, Any] (optional) - Metadata to associate with the trace. | Generate a TraceInfo object that will be used to group completions |
| `trace.record_output` | `output`: str - Output of the agent/trace. `project_id`: str - ID of the project to record to. `eval_results`: dict\[str, Any] (optional) - Code evaluation results. `test_run_info`: TestRunInfo (optional) | Record the output to a trace |
| `traces.update` | `project_id`: str - ID of the project. `session_id`: str - ID of the session. `trace_id`: str - ID of the trace. `output`: JSONValue (optional) - Updated output. `metadata`: dict (optional) - Custom metadata. `feedback`: dict (optional) - Customer feedback. `eval_results`: dict (optional) - Evaluation results. `test_run_info`: TestRunInfo (optional) - Test run information. | Update a trace after it has been recorded |
Traces are a more fine-grained way to group LLM interactions within a Session. A Trace can contain one or more completions and a Session can contain one or more Traces. Find a more detailed guide on how Sessions, Traces, and Completions fit together [here](/core-concepts/observability/sessions-traces-and-completions).
For a complete code example, see [Record Traces](/developer-resources/recipes/record-traces).
Traces are created off of an existing session object:
```python python theme={null}
input_question = "What color is the sky?"
# create or restore a session
session = fp_client.sessions.create()
# create the trace
trace_info = session.create_trace(
input=input_question,
agent_name="weather_agent",
custom_metadata={
"version": "1.0.8"
}
)
```
```typescript typescript theme={null}
const inputQuestion = "Why is the sky blue?"
// create or restore the session
const session = await fpClient.sessions.create();
// create the trace
const traceInfo = await session.createTrace({
input: inputQuestion,
// Optional parameters
agentName: "weather_agent",
customMetadata: {
version: "1.0.8"
}
});
// run series of LLM completions here
// from last llm call
const outputAnswer = "blue"
// record the trace
await traceInfo.recordOutput(
projectId,
outputAnswer, // LLM output (string)
// Code eval results (optional)
{
sentiment: 0.7,
validPath: true
}
);
```
```java java theme={null}
String inputQuestion = "Why is the sky blue?";
// 1. Create (or restore) a session
FreeplayClient fpClient = new FreeplayClient("YOUR_API_KEY", "https://app.freeplay.ai/api");
Session session = fpClient.sessions().create(projectId, "dev"); // specify environment if needed
// 2. Start a new trace with optional agentName and metadata
TraceInfo trace = session.createTrace(
inputQuestion,
"weather_agent", // agentName
Map.of("version", "1.0.8") // custom metadata
);
// 3. … run your LLM completions here …
String outputAnswer = "Blue"; // final LLM answer
// 4. Record the trace’s output along with evalResults
trace.recordOutput(
projectId,
outputAnswer,
// Eval results (optional)
Map.of(
"sentiment", 0.7,
"validPath", true
)
);
```
To tie a Completion to a given Trace you will pass the trace info in the record call
```python python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo, SessionInfo, TraceInfo
# create or restore a session
session = fp_client.sessions.create()
# create the trace
trace_info = session.create_trace(input=question)
# fetch prompt
# call LLM
# record with trace id
record_response = fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
session_info=session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=call_info,
trace_info=trace_info # Pass the trace info along
)
)
```
```typescript typescript theme={null}
import Freeplay, {getCallInfo, getSessionInfo} from "freeplay";
// create or restore the session
const session = await fpClient.sessions.create();
// create the trace
const traceInfo = await session.createTrace(userQuestion);
// fetch prompt
// call LLM
const completionResponse = await fpClient.recordings.create(
{
projectId,
allMessages: messages,
inputs: input_variables,
sessionInfo: session,
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
traceInfo: traceInfo // Pass the trace info
}
)
```
```java java theme={null}
// create or restore a session
Session session = fpClient.sessions().create();
// create the trace
TraceInfo trace = session.createTrace(inputQuestion);
// fetch prompt
// run llm completion
// record with the trace id
RecordInfo recordInfo = new RecordInfo(projectId, allMessages)
.inputs(variables)
.sessionInfo(sessionInfo)
.promptVersionInfo(formattedPrompt.getPromptInfo())
.callInfo(callInfo);
// add the trace info
recordInfo.traceInfo(trace);
```
For a complete working example, see [Record Traces](/developer-resources/recipes/record-traces). For agent-specific patterns, see [Agents](/practical-guides/agents).
Once you have recorded completions to the trace and are on the final output, you must close your trace in order to wrap the completions together, you can also optionally record eval results to the trace at this point:
```python python theme={null}
# record output to the trace
output_answer = "blue" # from the LLM
trace_info.record_output(
project_id=project_id,
output=output_answer,
eval_results={ # Optional trace eval logging
"sentiment": 0.7,
"valid_path": True,
}
)
```
```typescript typescript theme={null}
import Freeplay, { getCallInfo, getSessionInfo } from "freeplay";
// create or restore the session
const session = await fpClient.sessions.create();
// create the trace
const traceInfo = await session.createTrace(userQuestion);
// fetch prompt
// call LLM
const completionResponse = await fpClient.recordings.create({
projectId,
allMessages: messages,
inputs: input_variables,
sessionInfo: session,
promptInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end),
traceInfo: traceInfo, // Pass the trace info
});
```
```java java theme={null}
// create or restore a session
Session session = fpClient.sessions().create();
// create the trace
TraceInfo trace = session.createTrace(inputQuestion);
// fetch prompt
// run llm completion
// record with the trace id
RecordInfo recordInfo = new RecordInfo(
projectId,
allMessages,
variables,
session.getSessionInfo(),
formattedPrompt.getPromptInfo(),
callInfo,
);
// add the trace info
recordInfo.traceInfo(trace);
```
### Adding Tools to Traces
When building agents that use tools, tool calls are recorded as the output of an LLM call by default. You can also add explicit tool spans to provide more data about tool execution, including latency and other metadata. These are recorded as a Trace with `kind='tool'` and linked to the parent completion using `parent_id`.
For complete examples and code snippets, see [Tool Calls](/practical-guides/tools).
### Updating a Trace
Freeplay allows you to update a trace after it has been recorded. This is useful for adding evaluation results, customer feedback, metadata, or updating the output. To do this, you need the `project_id`, `session_id`, and `trace_id`. You must provide at least one of `output`, `metadata`, `feedback`, `eval_results`, or `test_run_info`.
```python python theme={null}
from freeplay.resources.traces import TraceUpdatePayload
# Update the trace with evaluation results and feedback
fp_client.traces.update(
TraceUpdatePayload(
project_id=project_id,
session_id=session.session_id,
trace_id=trace_info.trace_id,
output="updated output",
metadata={"version": "1.0.8"},
feedback={"freeplay_feedback": "positive", "satisfaction": True},
eval_results={"accuracy": 0.99},
)
)
```
```typescript typescript theme={null}
import { TraceUpdatePayload } from "freeplay";
// Update the trace with evaluation results and feedback
await fpClient.traces.update({
projectId,
sessionId: session.sessionId,
traceId: traceInfo.traceId,
output: "updated output",
metadata: { version: "1.0.8" },
feedback: { freeplay_feedback: "positive", satisfaction: true },
evalResults: { accuracy: 0.99 },
});
```
```java java theme={null}
import ai.freeplay.client.resources.traces.TraceUpdatePayload;
// Update the trace with evaluation results and feedback
fpClient.traces().update(
new TraceUpdatePayload(projectId, sessionId, String.valueOf(traceInfo.getTraceId()))
.output("updated output")
.metadata(Map.of("version", "1.0.8"))
.feedback(Map.of("freeplay_feedback", "positive"))
.evalResults(Map.of("accuracy", 0.99))
).get();
```
# About Freeplay
Source: https://docs.freeplay.ai/getting-started/freeplay-introduction
The ops platform for enterprise AI engineering teams
**Freeplay is the only platform your team needs to manage the end-to-end AI application development lifecycle.**
It provides an integrated workflow for improving your AI agents and other generative AI products. Engineers, data scientists, product managers, designers, and subject matter experts can all review production logs, curate datasets, experiment with changes, create and run evaluations, and deploy updates.
Here's a quick introduction.
## What Freeplay can do for you
* [AI Observability](#ai-observability)
* [Prompt management](#prompt-management)
* [Prompt playground](#prompt-playground)
* [Evaluations](#evaluations)
* [Testing](#testing)
* [Datasets](#datasets)
* [AI-powered features](#ai-powered-features)
* [Usage and costs](#usage-and-costs)
### AI Observability
Monitor your AI applications in real-time with powerful search, analytics for metrics like cost, latency, or custom evals you define, and automations to take quick action.
### Prompt management
Version prompts, models, and hyperparameters together, then log data against specific prompt versions for easy analysis. Optionally make Freeplay the source of truth for prompt and model configuration so non-engineers can deploy changes without code.
### Prompt playground
Experiment with prompt and model changes in a collaborative environment. Compare different versions side-by-side, test against saved datasets, and use AI-powered prompt optimization to speed up improvements.
### Evaluations
Define evaluators that measure quality both online (for production logs) and offline (for batch testing). Freeplay lets you define your own model-graded, code-based, and human evaluations.
### Testing
Run automated batch tests or evaluations at any time. Compare results between versions of individual prompts or complete agent workflows.
### Datasets
Curate test datasets from production logs, upload your own data, or create examples directly from the prompt editor.
### AI-powered features
Freeplay's AI agents help accelerate your product improvement workflow. Automated Review Insights turn human annotations and LLM judge scores into actionable themes. Prompt optimization uses your production data to generate improved prompts. And AI-assisted eval generation helps you write better evaluators faster.
### Usage and costs
Monitor and control LLM spend across all your model providers and environments. Track token usage, costs, and latency in real-time to optimize your AI application economics. See spend by project and by prompt.
Choose your path and start using Freeplay
# Integrate with Your Application
Source: https://docs.freeplay.ai/getting-started/integrate
Connect Freeplay to your AI application for observability, evaluations, and prompt management.
**Using a coding agent?** Point it to [docs.freeplay.ai/llms.txt](https://docs.freeplay.ai/llms.txt) for LLM-optimized documentation.
Integrating Freeplay with your application unlocks the full platform: production monitoring, dataset creation from real traffic, online and offline evaluations, and optional prompt and model deployment tools.
[Sign up for free](https://app.freeplay.ai/signup) to get started, or [reach out](https://freeplay.ai/demo) about self-hosting and enterprise options.
See the Glossary for definitions of key terms like sessions, traces, completions, and prompt templates.
## Choose your integration pattern
Before writing code, decide how you want to manage prompts.
### Pattern 1: Freeplay manages your prompts (recommended)
Freeplay becomes the source of truth for your prompt templates. Your application fetches prompts from Freeplay either at runtime, or as part of your build process (or both).
**Benefits:**
* Non-engineers can iterate on prompts and swap models without code changes
* Deploy prompt updates like feature flags and/or as part of your build process
* Automatic versioning and environment promotion
* Detailed observability with each log connected directly to a specific version (prompt and model configuration)
**Best for:** Teams that want to empower PMs, domain experts, or anyone else to iterate on prompts independently of code.
You can configure your Freeplay client to retrieve prompts from the server at runtime, or "bundle" them as part of your build process ([learn more about "prompt bundling"](/core-concepts/prompt-management/prompt-bundling)).
Many Freeplay customers retrieve prompts at runtime in lower-level environments like dev or staging to get the benefit of fast server-side experimentation, then use prompt bundling in production for tighter release management and zero latency.
### Pattern 2: Code manages your prompts
The source of truth for your prompts remains your codebase. You push prompt templates and model configurations to Freeplay to enable experimentation and organize your observability data.
**Benefits:**
* Prompts stay entirely in your code
* Use your existing code review process for prompt changes
* Full flexibility over prompt structure
* Sync prompts to Freeplay automatically via API in your CI/CD pipeline
**Best for:** Teams with complex prompt construction expectations, strict infrastructure-as-code requirements, and/or those who prefer prompt changes remain solely the domain of engineers.
With Pattern 2, you create prompt templates in Freeplay programmatically using the API. Use the `create_template_if_not_exists` parameter to sync prompts from your codebase:
```bash theme={null}
curl -X POST \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/prompt-templates/name/my-assistant/versions?create_template_if_not_exists=true" \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"template_messages": [
{"role": "system", "content": "You are a helpful assistant. The user'\''s name is {{user_name}}."},
{"role": "user", "content": "{{user_input}}"}
],
"provider": "openai",
"model": "gpt-4o",
"llm_parameters": {"temperature": 0.2, "max_tokens": 1024}
}'
```
See the full guide: [Code as source of truth](/core-concepts/prompt-management/managing-prompts#code-as-the-source-of-truth-for-prompts) and [Create prompt template version by name](/api-reference/prompt-templates/create-prompt-template-version-by-name).
### Optional: Start with observability only
You can begin by logging LLM calls to Freeplay without setting up prompt templates. This gets you started quickly, but **this is not recommended as an end state** since it limits core Freeplay features.
**What works without prompt templates:**
* Basic observability and search, including any evaluation values or other metadata you choose to log
* Manual agent dataset curation (trace level)
* Configuring auto-evaluations to run at the trace / agent level
**What requires prompt templates:**
* Experimenting in the Freeplay playground
* Creating prompt datasets from observability logs (requires `inputs` field / knowledge of variables in your prompts)
* Running evaluations that target prompt-level changes
* Searching by prompt template or version
* Running test runs against saved prompts
* Tracking prompt version performance over time
**When to use this approach:**
* You want to start logging immediately and add prompt templates later
* You're evaluating Freeplay before committing to a prompt management pattern
**Upgrade path**: Set up prompt templates as soon as possible to unlock the full set of Freeplay features. You can create templates either in the [Freeplay UI](/getting-started/start-in-ui) or [via the API](/core-concepts/prompt-management/managing-prompts#code-as-the-source-of-truth-for-prompts), then update your code to include `prompt_version_info` when recording.
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
from openai import OpenAI
import os
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api"
)
openai_client = OpenAI()
# Your existing prompt
messages = [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Hello!"}
]
# Make LLM call
response = openai_client.chat.completions.create(
model="gpt-4o",
messages=messages
)
# Record to Freeplay (no prompt template)
all_messages = messages + [
{"role": "assistant", "content": response.choices[0].message.content}
]
fp_client.recordings.create(
RecordPayload(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
all_messages=all_messages,
call_info=CallInfo(provider="openai", model="gpt-4o")
)
)
```
```typescript TypeScript theme={null}
import Freeplay from "freeplay";
import OpenAI from "openai";
const fpClient = new Freeplay({
freeplayApiKey: process.env.FREEPLAY_API_KEY,
baseUrl: "https://app.freeplay.ai/api"
});
const openaiClient = new OpenAI();
// Your existing prompt
const messages = [
{ role: "system", content: "You are helpful." },
{ role: "user", content: "Hello!" }
];
// Make LLM call
const response = await openaiClient.chat.completions.create({
model: "gpt-4o",
messages
});
// Record to Freeplay (no prompt template)
const allMessages = [...messages, response.choices[0].message];
await fpClient.recordings.create({
projectId: process.env.FREEPLAY_PROJECT_ID,
allMessages,
callInfo: { provider: "openai", model: "gpt-4o" }
});
```
## Get started
### Step 1: Install the SDK
You can integrate directly with your code using one of Freeplay's SDKs, or select from [integrations with common frameworks](#framework-integrations) outlined below.
```bash Python theme={null}
pip install freeplay
```
```bash TypeScript theme={null}
npm install freeplay
```
```groovy Java (Gradle) theme={null}
implementation 'ai.freeplay:freeplay-client-thin:VERSION'
```
```xml Java (Maven) theme={null}
ai.freeplayfreeplay-client-thinVERSION
```
#### Alternative: Framework integrations
If you're using a common AI framework, Freeplay provides native integrations that require separate packages from our standard SDKs. Learn more:
Instrument LangGraph agents with Freeplay
Add observability to Vercel AI applications
Integrate with Google's Agent Development Kit
Use standard OTel tracing with Freeplay
### Step 2: Create a prompt template
Create your first prompt template in the [UI](/getting-started/start-in-ui#1-create-your-first-prompt-template) or [via the API](/core-concepts/prompt-management/managing-prompts#code-as-the-source-of-truth-for-prompts).
Once you save a prompt, the **Integration** tab provides code snippets tailored to your template.
### Step 3: Integrate
Fetch prompts from Freeplay and log completions:
```python Python theme={null}
from freeplay import Freeplay
from openai import OpenAI
import os
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api"
)
openai_client = OpenAI()
project_id = os.getenv("FREEPLAY_PROJECT_ID")
## FETCH PROMPT ##
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="my-template",
environment="production",
variables={"user_input": "Hello, world!"}
)
## LLM CALL ##
response = openai_client.chat.completions.create(
model=formatted_prompt.model,
messages=formatted_prompt.llm_messages
)
## RECORD ##
all_messages = formatted_prompt.all_messages + [
{"role": "assistant", "content": response.choices[0].message.content}
]
fp_client.recordings.create(
project_id=project_id,
all_messages=all_messages,
prompt_version_info={
"prompt_template_version_id": formatted_prompt.prompt_template_version_id,
"environment": "production"
}
)
```
```typescript TypeScript theme={null}
import Freeplay from "freeplay";
import OpenAI from "openai";
const fpClient = new Freeplay({
freeplayApiKey: process.env.FREEPLAY_API_KEY,
baseUrl: "https://app.freeplay.ai/api"
});
const openaiClient = new OpenAI();
const projectId = process.env.FREEPLAY_PROJECT_ID;
/* FETCH PROMPT */
const formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "my-template",
environment: "production",
variables: { userInput: "Hello, world!" }
});
/* LLM CALL */
const response = await openaiClient.chat.completions.create({
model: formattedPrompt.model,
messages: formattedPrompt.llmMessages
});
/* RECORD */
const allMessages = [
...formattedPrompt.allMessages,
response.choices[0].message
];
await fpClient.recordings.create({
projectId,
allMessages,
promptVersionInfo: {
promptTemplateVersionId: formattedPrompt.promptTemplateVersionId,
environment: "production"
}
});
```
```java Java theme={null}
import ai.freeplay.client.thin.Freeplay;
import ai.freeplay.client.thin.resources.prompts.FormattedPrompt;
import com.openai.client.OpenAIClient;
import com.openai.models.ChatCompletion;
String apiKey = System.getenv("FREEPLAY_API_KEY");
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
Freeplay fpClient = new Freeplay(
Freeplay.Config()
.freeplayAPIKey(apiKey)
.baseUrl("https://app.freeplay.ai/api")
);
OpenAIClient openaiClient = OpenAIClient.builder().build();
/* FETCH PROMPT */
Map variables = Map.of("userInput", "Hello, world!");
FormattedPrompt formattedPrompt = fpClient.prompts().getFormatted(
projectId,
"my-template",
"production",
variables
).get();
/* LLM CALL */
ChatCompletion response = openaiClient.chat().completions().create(
ChatCompletionCreateParams.builder()
.model(formattedPrompt.getModel())
.messages(formattedPrompt.getLlmMessages())
.build()
);
/* RECORD */
List
Define prompts in your code and log to Freeplay:
```python Python theme={null}
from freeplay import Freeplay, RecordPayload, CallInfo
from openai import OpenAI
import os
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base="https://app.freeplay.ai/api"
)
openai_client = OpenAI()
project_id = os.getenv("FREEPLAY_PROJECT_ID")
## YOUR PROMPT (defined in code) ##
prompt_vars = {"user_name": "Alice", "topic": "weather"}
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": f"Hi {prompt_vars['user_name']}, tell me about {prompt_vars['topic']}."}
]
## LLM CALL ##
model = "gpt-4.1-mini"
response = openai_client.chat.completions.create(
model=model,
messages=messages
)
## RECORD ##
all_messages = messages + [
{"role": "assistant", "content": response.choices[0].message.content}
]
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars, # Variables enable dataset creation
call_info=CallInfo(provider="openai", model=model)
)
)
```
```typescript TypeScript theme={null}
import Freeplay from "freeplay";
import OpenAI from "openai";
const fpClient = new Freeplay({
freeplayApiKey: process.env.FREEPLAY_API_KEY,
baseUrl: "https://app.freeplay.ai/api"
});
const openaiClient = new OpenAI();
const projectId = process.env.FREEPLAY_PROJECT_ID;
/* YOUR PROMPT (defined in code) */
const promptVars = { userName: "Alice", topic: "weather" };
const messages = [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: `Hi ${promptVars.userName}, tell me about ${promptVars.topic}.` }
];
/* LLM CALL */
const model = "gpt-4.1-mini";
const response = await openaiClient.chat.completions.create({
model,
messages
});
/* RECORD */
const allMessages = [...messages, response.choices[0].message];
await fpClient.recordings.create({
projectId,
allMessages,
inputs: promptVars, // Variables enable dataset creation
callInfo: { provider: "openai", model }
});
```
```java Java theme={null}
import ai.freeplay.client.thin.Freeplay;
import com.openai.client.OpenAIClient;
import com.openai.models.ChatCompletion;
String apiKey = System.getenv("FREEPLAY_API_KEY");
String projectId = System.getenv("FREEPLAY_PROJECT_ID");
Freeplay fpClient = new Freeplay(
Freeplay.Config()
.freeplayAPIKey(apiKey)
.baseUrl("https://app.freeplay.ai/api")
);
OpenAIClient openaiClient = OpenAIClient.builder().build();
/* YOUR PROMPT (defined in code) */
Map promptVars = Map.of("userName", "Alice", "topic", "weather");
List
To fully unlock Freeplay features like searching by prompt template or version in Observability, sync the prompts from your code to Freeplay using the API. See [Code as source of truth](/core-concepts/prompt-management/managing-prompts#code-as-the-source-of-truth-for-prompts) for the complete workflow and [Create prompt template version by name](/api-reference/prompt-templates/create-prompt-template-version-by-name) for endpoint details.
## Next steps
Your integration is complete. Sessions will appear in the [Observability dashboard](/core-concepts/observability/observability-dashboard) as you log data from your application.
With prompts and observability configured, you can now set up evaluations and run tests.
Learn more about evaluation types and alignment
Execute batch tests to compare prompt versions and measure quality
# Overview
Source: https://docs.freeplay.ai/getting-started/overview
Choose your path to get started with Freeplay
There are two ways to start using Freeplay, depending on your role and goals.
**Best for**: Product managers, domain experts, or developers who want to explore Freeplay before integrating.
Create prompts, build datasets, set up evaluations, and run tests entirely in the Freeplay UI. No code required.
**Best for**: Developers ready to connect Freeplay to an existing application.
Add observability to your AI application and unlock the full platform including automated evals, prompt management, and more.
## Which path should you choose?
### Start in the UI
Choose this path if you want to use Freeplay before you've built an app, or before you're ready to integrate. You can:
* Import prompts and begin to test changes against different models
* Define evaluations and compare different prompts and models
* Build test datasets and run your own custom evaluations before writing any code
* Learn Freeplay's capabilities now, then integrate working prompt templates and evals with your code when ready
### Integrate with your app
Choose this path if you've already built an app or agent and want to:
* Capture, view, and score agent traces or individual LLM completions when your application runs
* Run evaluations on live data to generate production metrics
* Coordinate manual review workflows across your team
* Create datasets from actual user interactions or production traces
* Start managing prompt and model configurations in Freeplay
* Create a testing harness to run your own custom evals any time you change your code
## Get your Freeplay account
Create your Freeplay account and start building in minutes
Interested in self-hosting or enterprise features? Schedule a call with our team
# Start in the UI
Source: https://docs.freeplay.ai/getting-started/start-in-ui
Create prompts, datasets, evaluations, and tests without writing code
You can accomplish a lot in Freeplay without integrating any code. This guide walks you through the core workflows you can complete entirely in the UI.
**Need a Freeplay account?** [Sign up for free](https://app.freeplay.ai/signup) to get started, or [reach out](https://freeplay.ai/demo) about self-hosting and enterprise options.
## What you can do
* **Create and iterate on prompts** in the playground
* **Build test datasets** manually or by uploading CSV/JSONL files
* **Set up evaluations** including model-graded and human label values
* **Run tests** to compare prompt and model changes quantitatively
When you're ready to monitor production traffic or create datasets from real user interactions, see the [integration guide](/getting-started/integrate).
## 1. Create your first prompt template
From your project, click **Create prompt template** to open the playground.
### Select your model
Choose from available default providers like OpenAI, Anthropic, or Google. When you're ready, you'll want to [configure your own API keys](/account-setup/project#configure-models).
### Add messages and variables
Freeplay prompt templates use **messages** (static content) and **variables** (dynamic inputs using `{{variable_name}}` syntax). This separation enables:
* Rapidly building datasets from actual logs
* Setting up evaluations that reference or compare specific input variables
* Easy batch testing across different inputs or test cases
#### Message types
* **System message**: Sets the AI's behavior and personality. This is your base instruction that defines how the AI should act.
* **User message**: Represents input from your end user. Use variables to make it dynamic: `Create an album name for {{artist}}`
* **Assistant message**: Pre-filled AI responses that can serve different purposes, e.g. few-shot examples.
* **History**: A special message type in Freeplay that represents conversation history in multi-turn conversations or maintains context in agent traces.
### Configure advanced settings
Beyond messages, you can set model-specific hyperparameters like temperature, max tokens, and other model parameters. You can also add tools for function calling or enable structured outputs.
### Test in the playground
Before saving, test your prompt across a range of examples by loading saved datasets or manually entering values for each variable and clicking **Run**. Outputs generated in the playground can be saved to start building your first dataset.
### Save your prompt template
Once you're satisfied, click **Save** and name your prompt template. Optionally add a version name and description to help your team understand what changed.
Each version of the prompt template that you save includes the prompt text and variables, the model and provider selected, and any hyperparameters you set.
[Learn more about prompt templates](/core-concepts/prompt-management/managing-prompts)
## 2. Create a dataset
Datasets power your evaluations and test runs. You can:
* **Save from the playground**: Save examples manually as you iterate in the playground
* **Upload data**: Import CSV or JSONL files with test cases
* **Add manually**: Create individual examples in the datasets UI
Each dataset row can include expected outputs for reference-based evaluations. You can also get started with just sample inputs.
[Learn more about datasets](/core-concepts/datasets/datasets)
## 3. Set up evaluations
Define how you'll measure quality. Freeplay supports:
* **Model-graded evaluations**: Use an LLM judge to score outputs
* **Auto-categorization**: Automatically classify outputs against criteria
* **Code evaluations**: Custom logic (requires integration)
* **Human labels**: Manual review by team members
[Learn more about evaluations](/core-concepts/evaluations/evaluations)
## 4. Run batch tests
Once you have a prompt, dataset, and evaluations configured:
1. Navigate to your prompt template
2. Click **Run test**
3. Select your dataset and the evaluators you want to run
4. Compare results across versions
Tests give you quantitative data to decide which prompt and model combinations perform best. You can also dig into row-level data for any test result to see exactly how a given set of inputs perform.
[Learn more about test runs](/core-concepts/test-runs/test-runs)
## Next steps
Connect Freeplay to your application for production monitoring
Learn more about evaluation types and alignment
# Why Freeplay?
Source: https://docs.freeplay.ai/getting-started/why-freeplay
Why AI engineering teams choose Freeplay as their ops platform for evals, observability, testing and experimentation
Building AI products is fundamentally different from building traditional software. Success depends on how fast your team can iterate: identifying what's working, experimenting with improvements, and validating changes before they reach users. The teams that execute this cycle fastest ship better products.
Freeplay is built from the ground up to accelerate this iteration cycle. Where other tools can offer a collection of developer utilities and leave teams to figure out how the pieces connect, Freeplay provides an opinionated workflow that guides your team through a continuous feedback loop from production insights to tested improvements.
## A connected workflow
AI engineering teams quickly discover they need common components to support their ops workflow -- like logging, playgrounds, and evaluation runners. But these often feel disconnected from each other. Each solves a narrow problem without consideration for what comes before or after.
Freeplay takes an integrated approach. Every feature is designed to feed into the next step of your iteration cycle:
* **Prompt templates and named agents separate structure from data.** Freeplay distinguishes between the static parts of your prompts and the variables populated at runtime. Each instance of a trace for a specific agent or sub-agent gets named and grouped together. This structure makes it seamless to turn production logs into replayable dataset rows for testing.
* **Datasets enforce compatibility.** Dataset schemas match prompt templates or specific agents, so you always know whether a dataset will in a given test scenario. No more time lost reformatting test data.
* **Evaluations reference your data structure.** When writing LLM judges or code-based evaluators, you can target or compare specific input variables (not just an entire interpolated input blob). This precision leads to more meaningful quality signals.
* **Production traces become test cases.** Annotated examples from production flow directly into datasets. Failures become regression tests. The system is designed to turn usage into better testing data.
This connected design provides the foundation for building a **data flywheel**: where each iteration strengthens not just your core prompts and agents, but also your datasets, evaluation criteria, and testing infrastructure. Over time, those elements compound into a closed loop for continuous improvement.
## True cross-functional collaboration
AI product development works best when engineers, product managers, designers, and domain experts work together. The people closest to customer or business problems often have the clearest sense of what "good" looks like, but most AI engineering tools relegate them to spectators who can only contribute when closely supported by an engineer.
Freeplay changes this dynamic, so that each team member can contribute their full expertise. Non-engineers can:
* Create and iterate on prompts, models and tool definitions in the playground
* Build test datasets manually or from production logs
* Write and refine LLM judges for custom evaluation metrics
* Run tests and evaluations to compare prompt and model changes
* Review agent traces and annotate quality issues
All of those can be completed without touching code. Engineers stay in control of orchestration and deployment, while domain expertise flows directly into the product.
[Learn what you can do in the UI →](/getting-started/start-in-ui)
## AI that accelerates your workflow
The future of AI engineering involves AI agents working alongside human teams. Freeplay applies AI at specific points in the iteration cycle where it adds the most value, for example:
* **Eval generation** helps you write better LLM judges faster, automatically adapting to your prompt structure and data
* **Prompt optimization** uses your production data -- evaluation results, user feedback, and human annotations -- to generate improved prompt versions, optimized for specific models
* **Review insights** analyze patterns across human notes and LLM judge reasoning to surface actionable themes and root causes
These features are designed to amplify your team's expertise, not replace it. Humans define the goals and provide nuanced feedback, and AI accelerates the path to meeting those goals.
## Built for enterprise teams
Freeplay truly serves the needs of enterprise product development teams: organizations with strong software engineering foundations applying AI to complex business problems. We focus on teams that need:
* **Production-grade infrastructure** that works in any cloud and scales with usage, providing instant search over terrabytes of logs and traces
* **Framework flexibility** to work with your existing stack, whether you write your own custom code or use popular agent frameworks
* **Security and compliance** controls for enterprise requirements, including support for multi-region deployments to support strict data domicile requirements
* **Premium support** to help teams without prior AI engineering experience build solid evaluations and test harnesses and adopt best practices
The data asset you build in Freeplay -- your logs, evaluations, and datasets -- belongs to you. It gives you full flexibility to switch models, providers, or libraries as the ecosystem evolves.
## Get started
Create prompts, datasets, and evaluations without writing code
Connect Freeplay to your existing AI application
# Introduction
Source: https://docs.freeplay.ai/openapi/introduction
Observe, evaluate, and iterate toward great AI applications
The Freeplay REST API enables you to integrate observability and prompt management into your AI applications.
Access the raw OpenAPI 3.0 specification for use with code generators, API clients, or LLM tools.
**Using a coding agent?** Point it to [docs.freeplay.ai/llms.txt](https://docs.freeplay.ai/llms.txt) for LLM-optimized documentation.
## Base URL
Your API root is your Freeplay instance URL plus `/api/v2/`:
```
https://app.freeplay.ai/api/v2
```
For private deployments, use your custom domain (e.g., `https://acme.freeplay.ai/api/v2`).
## Authentication
Freeplay authenticates requests using API keys, which you can manage at `https://app.freeplay.ai/settings/api-access`.
Include your API key in the `Authorization` header:
```
Authorization: Bearer {freeplay_api_key}
```
## Error Handling
Freeplay uses standard HTTP status codes:
| Code | Meaning |
| ---- | ------------------------------------------- |
| 200 | Success |
| 400 | Bad request — check your request parameters |
| 401 | Unauthorized — invalid or missing API key |
| 500 | Server error — safe to retry |
Errors in the 400-499 range are client errors and won't resolve with retries.
For 500 errors, retry up to 3 times with at least a 5-second delay between
attempts.
# Limits
Source: https://docs.freeplay.ai/openapi/limits
API limits and error messages to help you debug integration issues faster.
Freeplay APIs include limits and error messages to help you know when requests fall outside expected bounds and debug issues faster. If a request exceeds a limit, you'll receive a clear error message so you can adjust accordingly.
| Resource | Limit | Exceeded behavior |
| -------------------------------------------------------------------- | --------------- | ------------------------------------------------ |
| Completions/traces per session (hard cap) | 10,000 | HTTP 400 |
| Completions/traces per session (indexed) | 1,000 | Additional records are stored but not searchable |
| Unique keys per structured field (structured output, metadata, etc.) | 1,000 per field | HTTP 400 |
| Record payload size (per individual completion or trace) | 10 MB | HTTP 413 |
| Test cases per dataset | 3,000 | HTTP 400 |
Please reach out to [support@freeplay.ai](mailto:support@freeplay.ai) if you any of these are limiting for your use case.
# Search API Operators
Source: https://docs.freeplay.ai/openapi/search-api-operators
Filter operators and query syntax for the Search API endpoints
Use Search endpoints to build complex queries and retrieve your data via API.
The operators below can be used to build filters that help you find relevant sessions, traces, and completions based on cost, latency, metadata, online evaluation results, human labels or notes, and more.
This guide supplements the Search API endpoints from the OpenAPI specification. While the spec defines the endpoint structure and request/response schemas, this page provides detailed documentation on all available filter operators and how to construct complex queries.
## Endpoints
| Endpoint | Description | API Reference |
| -------------------------- | ------------------ | ------------------------------------------------------------------- |
| `POST /search/sessions` | Search sessions | [Search Sessions](/developer-resources/api-reference#search-api) |
| `POST /search/traces` | Search traces | [Search Traces](/developer-resources/api-reference#search-api) |
| `POST /search/completions` | Search completions | [Search Completions](/developer-resources/api-reference#search-api) |
All endpoints support pagination via `page` and `page_size` query parameters.
```bash theme={null}
curl -X POST \
-H "Authorization: Bearer $FREEPLAY_API_KEY" \
-H "Content-Type: application/json" \
"https://app.freeplay.ai/api/v2/projects/$PROJECT_ID/search/completions?page=1&page_size=20" \
-d '{"filters": {"field": "cost", "op": "gte", "value": 0.01}}'
```
## Filter Operators
| Operator | Description |
| ---------- | --------------------- |
| `eq` | Equals |
| `lt` | Less than |
| `gt` | Greater than |
| `lte` | Less than or equal |
| `gte` | Greater than or equal |
| `contains` | Contains substring |
| `between` | Within numeric range |
## Available Filters
The table below lists all available filter fields, their supported operators, and example values. Fields ending in `.*` support filtering on nested JSON properties.
| Field | Supported Operators | Example Value |
| ---------------------------------------- | ------------------------------------------ | --------------------------- |
| `cost` | `eq`, `lt`, `gt`, `lte`, `gte` | `0.003` |
| `latency` | `eq`, `lt`, `gt`, `lte`, `gte` | `8` |
| `start_time` | `eq`, `lt`, `gt`, `lte`, `gte` | `"2024-06-01 00:00:00"` |
| `environment` | `eq` | `"staging"` |
| `prompt_template` | `eq` | `"my-prompt"` |
| `prompt_template_id` | `eq` | `"uuid..."` |
| `completion_id` | `eq` | `"uuid..."` |
| `session_id` | `eq` | `"uuid..."` |
| `test_run_id` | `eq` | `"uuid..."` |
| `review_queue_id` | `eq` | `"uuid..."` |
| `model` | `eq` | `"gpt-4o"` |
| `provider` | `eq` | `"openai"` |
| `review_status` | `eq` | `"review_complete"` |
| `agent_name` | `eq` | `"support-agent"` |
| `trace_agent_name` | `eq` | `"my-agent"` |
| `api_key` | `eq` | `"production-key"` |
| `assignee` | `eq` | `"user@example.com"` |
| `insight_name` | `eq` | `"Response Quality Issues"` |
| `completion_output` | `contains` | `"weather"` |
| `completion_inputs.*` | `contains` | `"topic": "weather"` |
| `completion_feedback.*` | `contains` | `"rating": "positive"` |
| `session_custom_metadata.*` | `contains` | `"user_type": "premium"` |
| `trace_custom_metadata.*` | `contains` | `"workflow": "onboarding"` |
| `trace_input.*` | `contains` | `"query": "weather"` |
| `trace_output.*` | `contains` | `"response": "sunny"` |
| `trace_feedback.*` | `contains` | `"rating": "positive"` |
| `completion_evaluation_results.*` | `eq` | `"Response Quality": "4"` |
| `completion_client_evaluation_results.*` | `eq` | `"score": "85"` |
| `trace_evaluation_results.*` | `eq`, `gt`, `lt`, `gte`, `lte`, `contains` | `"Quality Score": 5` |
| `trace_client_eval_results.*` | `eq`, `contains` | `"confidence_score": 0.95` |
| `evaluation_notes.content` | `contains` | `"needs review"` |
| `evaluation_notes.author` | `eq` | `"user@example.com"` |
| `evaluation_notes.created_at` | `gt`, `lt`, `gte`, `lte` | `"2024-06-01 00:00:00"` |
## Compound Filters
Combine multiple filters using logical operators (`and`, `or`, `not`) to build complex queries. These operators can be nested to any depth.
### Using `and`
All conditions must be true:
```json theme={null}
{
"filters": {
"and": [
{"field": "cost", "op": "gte", "value": 0.001},
{"field": "model", "op": "eq", "value": "gpt-4o"}
]
}
}
```
### Using `or`
At least one condition must be true:
```json theme={null}
{
"filters": {
"or": [
{"field": "model", "op": "eq", "value": "gpt-4o"},
{"field": "model", "op": "eq", "value": "claude-3-opus"}
]
}
}
```
### Using `not`
Negates a condition:
```json theme={null}
{
"filters": {
"not": {"field": "environment", "op": "eq", "value": "prod"}
}
}
```
### Combined Example
Find expensive completions using either GPT-4o or Claude 3 Opus, excluding production:
```json theme={null}
{
"filters": {
"and": [
{"field": "cost", "op": "gte", "value": 0.001},
{
"or": [
{"field": "model", "op": "eq", "value": "gpt-4o"},
{"field": "model", "op": "eq", "value": "claude-3-opus"}
]
},
{
"not": {"field": "environment", "op": "eq", "value": "prod"}
}
]
}
}
```
# Agents in Freeplay
Source: https://docs.freeplay.ai/practical-guides/agents
Overview of how to build, monitor, and improve AI agents using Freeplay.
# Agents in Freeplay
## How We Use The Word Agent
The term “Agent” is widely used and can mean different things. For our purposes here, we mean *any* process in your application that makes use of multiple LLM calls to generate a single response. These may include tool usage or self-directed reasoning processes, but they don’t have to.
Everything below applies to both “workflows” and “agents” as defined in Anthropic’s helpful guide to [“Building Effective Agents.”](https://www.anthropic.com/engineering/building-effective-agents)
## Overview
Many LLM applications now go beyond simple completions to include multi-step workflows where an LLM makes decisions, uses tools, and follows complex reasoning processes to accomplish a task. These agent-based systems require specialized tooling to effectively monitor, test, and improve.
Freeplay provides robust support to build and improve AI agents — allowing you to observe, evaluate, test changes, and iterate on agent performance, no matter how you manage orchestration. To get started, you'll need to:
* Configure [prompt management](/core-concepts/prompt-management/managing-prompts) in Freeplay, so that any prompts included in your agent can be tracked and versioned
* [Record traces](/freeplay-sdk/traces) to Freeplay any time your agent runs (including any tool calls, additional metadata, customer feedback, and any evals you calculate in your code)
This guide will walk you through the details of how to effectively use Freeplay for your agent workflows.
### Key Benefits
* **Agent Observability:** Monitor and observe each of the steps, decisions, and outputs in your agent workflows to identify areas for improvement
* **Automated Testing & Structured Experimentation**: Test agent components individually or as complete workflows — with support to manage evals and testing datasets at both the individual component and full agent level
* **Advanced Evaluations**: Apply evaluations to any individual agent steps and/or overall agent performance
* **Collaborative Debugging & Iteration**: Easily share agent performance data and logs in ways that are easy to interpret for your whole team, in support of collaborative problem-solving
* **Performance Insights**: Quantify your agent’s performance on metrics you define, and report out findings to your wider organization
## How Agents Work In Freeplay
Freeplay depends on the concept of **traces** to support agent workflows.
A Freeplay trace represents a logical grouping of LLM completions, tool calls, etc. that form a single agent task or workflow. Each trace can contain multiple completions, tool calls/results, and additional metadata about the agent — along with evaluation results and user feedback.
### Relationship between Sessions, Traces, and Completions
Understanding the hierarchy of data organization in Freeplay is essential for effectively tracking agent workflows. In short:
* **Completions:** Individual LLM calls made up of a prompt and a response or output from a model.
* **Traces:** Optionally used to group related completions and tool calls, e.g. when multiple completions are used to generate a single chat turn or an agent flow.
* **Sessions:** The container for all completions and traces that make up a single customer interaction or single agent run.
* Note: These can be 1:1 with completions for a simple feature that just uses one prompt, or they can be very large at times, e.g. an entire conversation thread between a single user and a chatbot over multiple hours where each turn invokes an agentic workflow.
For agent workflows, you'll typically have:
1. One session per user interaction or single run of your agent (unless you're [building a multi-turn chatbot](/practical-guides/multi-turn-chat-support))
2. One or more traces per agent or multi-step workflow
3. Multiple completions within each trace
For more detailed information on this data hierarchy, please refer to our guide on [Sessions, Traces and Completions](/core-concepts/observability/sessions-traces-and-completions).
## Logging Data to Agents
When logging agent activity, include not just the inputs and outputs, but also metadata about the agent version, capabilities, and any tool calls made during the process. The following dives into the details of how to log this data to Freeplay:
### Agent Creation
To track your agents with Freeplay, you'll create traces and record agent interactions. When creating a trace for an agent, you can include:
* `agent_name` (string): The name of the agent being used, which gets used elsewhere in Freeplay for evals, dataset compatibility, etc.
* `custom_metadata` (key-value pairs): Additional metadata you choose to record about the agent
```python python theme={null}
# Create an example trace for an agent
trace = session.create_trace(
input="customer_inquiry: What's my account balance?",
agent_name="Financial Assistant",
custom_metadata={
"agent_version": "2.1.3",
"agent_framework": "multi-modal agent",
"agent_capabilities": "banking,accounts,transfers"
}
)
```
```typescript typescript theme={null}
import { CustomMetadata } from "freeplay";
// Example inputs
const userQuestion = "What is the weather in London?";
const agentName = "WeatherAgent";
const customMetadata: CustomMetadata = {
location: "London",
requestedAt: new Date().toISOString(),
};
// Creating the trace
const traceInfo = await session.createTrace({
input: userQuestion,
agentName,
customMetadata,
eval_result,
});
```
### Recording Intermediate Completions & Tool calls within an Agent
For each step in your agent's workflow, you'll want to record the LLM completions and any tool calls made. To record this information to a trace, pass the `trace_info` to the `RecordPayload` function along with any completion or tool call/results and they will be tied to the trace.
#### Two approaches to logging tool calls
Freeplay supports two approaches for recording tool calls:
| Approach | Best for | How it works |
| ---------------------------- | ------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Default (in completions)** | Simpler integrations, standard tool-calling patterns | Tool calls are recorded as part of the message history. The tool call appears in the LLM's output message, and the tool result appears as an input message to the next LLM call. When you call `formatted_prompt.all_messages()`, tool calls and results are concatenated into the `history` object alongside other messages. |
| **Explicit tool spans** | Complex agent debugging, tool execution timing, separate tool result visibility | Create separate traces with `kind='tool'` for each tool execution. This renders tool arguments and results as independent spans in the trace view, in addition to their presence in the message history. |
Most integrations should start with the default approach. Use explicit tool spans when you need to debug tool execution timing, analyze tool performance separately, or surface tool behavior more prominently in observability.
For detailed implementation of both approaches, see [Tool Calls](/practical-guides/tools#logging-tool-calls-to-freeplay).
See our full tool calling examples for [OpenAI](/developer-resources/recipes/using-tools-with-openai) and [Anthropic](/developer-resources/recipes/using-tools-with-anthropic). The code below shows how to record tool calls and intermediate completions:
```python python theme={null}
######
# LLM Call with tool calls in the response
######
# Handle tool calls if present
if isinstance(completion.content, list):
for block in completion.content:
if isinstance(block, ToolUseBlock) and block.name == "weather_of_location":
temperature = get_temperature(
block.input["location"]
) # Capture the tool response in the right format
tool_response_message = {
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": block.id,
"content": str(temperature),
}
]
}
messages.append(tool_response_message)
# Record a tool call using RecordPayload
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted_prompt.prompt_info,
start,
end
),
tool_schema=formatted_prompt.tool_schema,
trace_info=trace_info
)
)
```
### Recording LLM Outputs and Client-Side Evaluations
When recording agent outputs, you can also include any evaluation results calculated in your code ([details here](/core-concepts/evaluations/code-evaluations)):
* `eval_results` (dict): Dictionary of evaluation metrics for the trace (can include numeric, string, or boolean values)
```python python theme={null}
# Record trace output with evaluations
trace.record_output(
project_id=PROJECT_ID,
output=agent_response,
eval_results={
"task_completion": True,
"reasoning_quality": 0.85,
"response_accuracy": "high",
"hallucination_detected": False
}
)
```
```typescript typescript theme={null}
// Typing information
type CustomFeedback = {
freeplay_feedback: "positive" | "neutral" | "negative";
is_helpful: boolean;
};
// Variable prep
type TestRunInfo = {
runId: string;
runName?: string;
};
const botResponseText = botResponse.llmResponseText;
const evalResults = {
is_factual: true,
helpfulness_score: 0.9,
};
const testRunInfo: TestRunInfo = {
runId: "test-run-123",
runName: "RegressionSuite-May",
};
// Record the model output along with optional eval results and test run info
await traceInfo.recordOutput(
projectId,
botResponseText,
evalResults,
testRunInfo
);
```
### Updating Agents with Customer Feedback
You can add customer or system feedback to traces, which will be treated as a special class of metadata in Freeplay.
```python python theme={null}
# Update trace with feedback
client.customer_feedback.update_trace(
project_id=PROJECT_ID,
trace_id=trace.trace_id,
feedback={
"helpfulness": 9.2,
"relevance": "high",
"satisfied_user": True,
"valid_tool_use": True,
"guard_rails": False
}
)
```
```typescript typescript theme={null}
// Typing information
const feedback: Record = {
freeplay_feedback: {
freeplay_feedback: "positive",
is_helpful: true,
},
};
// Update the trace with customer feedback
await fpClient.customerFeedback.updateTrace({
projectId,
traceId: traceInfo.traceId,
customerFeedback: feedback,
});
```
## Agent Evals
Monitor and improve your agent's performance with Freeplay's evaluation system. Agent Evals help you track whether your agents are delivering the quality and accuracy you expect in production.
### What Agent Evals Do
Agent Evals continuously assess your agent's outputs using both AI models and human reviewers. Think of them as quality checks that run alongside your production system, alerting you when performance drifts from your standards. You can evaluate various aspects like response usefulness, clarity, accuracy, and whether the agent is actually solving the problem it was given.
### Setting Up Your Agent Eval
Getting started is straightforward. Once you have at least one recorded trace from your agent, you can create an eval by:
1. Define your goal - What aspect of performance matters most? (e.g., "Is the response helpful?" or "Did the agent extract the correct information?")
2. Set clear criteria - Describe what success looks like for this evaluation
3. Configure monitoring - Choose your evaluation model and set how often to sample live production data
4. Deploy and track - Save your eval and start monitoring your agent's real-world performance
The real power comes from running these evaluations continuously in production. You'll spot issues early and maintain consistent quality as your system scales and evolves.
## Agent Datasets & Testing
### Datasets
Agent datasets enable end-to-end testing of your agent workflows, allowing you to assess the complete behavior of your system under real or representative conditions. These datasets serve as a foundation for understanding how changes—such as prompt revisions, tool updates, or orchestration logic adjustments—affect the overall output of your agent.
In Freeplay, agent datasets are powered by trace-level logging. By assigning an Agent Name to your traces, you create a logical grouping of all related agent runs. This grouping allows you to filter and review traces in the observability dashboard and assemble them into datasets for future testing and evaluation.
You can create an agent dataset in two ways:
1. From the observability dashboard by selecting traces with the same Agent Name.
2. Directly from the trace view by saving specific traces to a dataset.
### Testing
Once you’ve assembled an agent dataset, you can use it to run structured tests against your agent. These tests provide visibility into system performance, surfacing both top-level metrics and step-level details that inform iteration and deployment decisions.
Running tests on agent datasets allows you to apply evaluations at both the trace and prompt level and understand how changes impact overall system behavior.
The example below illustrates a completed test run in Freeplay. High-level agent evaluations are shown at the top, while granular prompt-level evaluations are displayed below. This layered view enables you to assess both the success of the full agent and the contributions of each step.
To execute these tests programmatically, you can use the Freeplay SDK in the same way you would initiate a standard test run. See the [SDK documentation](/freeplay-sdk#trace-test-runs) for full implementation details.
## FAQ
### How do I update traces in my current Freeplay implementation?
If you're already using Freeplay, you'll need to:
1. Update to the latest SDK version
2. Modify your trace creation code to include `agent_name` and `custom_metadata`
3. Update your recording logic to include evaluations and/or customer feedback as needed
### Do I need to use traces to work with agents?
While not strictly required, using the `agent_name` value with traces provides significant benefits for building agents:
* Better visibility into multi-step processes
* Ability to evaluate agents from end to end
* Ability to run and organize tests at the agent level, separate from individual components
* More granular performance metrics
* Enhanced debugging capabilities
### How does a trace map to an agent?
The relationship between traces and agents is flexible and ultimately up to you:
* One trace can represent one complete agent task
* Multiple traces can represent different aspects of a complex agent
Choose the approach that best represents your agent's logical workflow.
## Best Practices
1. **Naming Convention**: Use consistent agent naming to make searching and analysis easier
2. **Metadata Strategy**: Define a minimum standard set of metadata fields for your agents up front. (You can always add more later too.)
3. **Granular Evaluations**: Create evaluations that target specific agent components *and* end-to-end behaviors
4. **Representative Datasets**: Build datasets that cover the range of expected agent tasks
5. **Regular Testing**: Build batch testing with evals into your agent development workflow to catch issues and regressions early
## Next Steps
Ready to get started with agents in Freeplay?
1. Update your Freeplay SDK to the latest version
2. Review our code examples for implementing agent trace logging
3. Create your first agent dataset and evaluations (where the dataset consists of inputs and outputs for the entire end-to-end agent behavior)
4. Set up dashboards to monitor agent performance
For more detailed implementation guidance, contact our support team or schedule a consultation with our forward deployed engineering team.
# Voice-Enabled AI with Pipecat, Twilio, and Freeplay
Source: https://docs.freeplay.ai/practical-guides/build-voice-enabled-ai-applications-with-pipecat-twilio-and-freeplay
Build and monitor voice-enabled AI applications using Pipecat, Twilio, and Freeplay.
## Introduction
Voice-enabled AI applications present unique challenges when it comes to testing, monitoring, and iterating on your prompts and models. This guide demonstrates how Freeplay's observability and prompt management tools can support your development workflow when building voice applications.
In this example, we show how to use Freeplay together with Pipecat and Twilio.
## What is Pipecat?
[Pipecat](https://docs.pipecat.ai/getting-started/overview) is a powerful open source framework for building voice-enabled, real-time, multimodal AI applications.
When paired with [Twilio](https://www.twilio.com/en-us) for real-time voice over the phone, Pipecat enables teams to quickly build audio-based agentic systems that combine both user and bot audio with LLM interactions.
This combination creates a strong foundation for the core application, but building a high-quality generative AI product also requires robust monitoring, evaluation, and continuous experimentation. This is where Freeplay helps.
## Using Freeplay for Rapid Iteration and Observability
When it comes to monitoring and improving a voice agent, teams often struggle with:
* **Multi-modal Observability**: Tracking and analyzing model inputs and outputs across different data types (audio, text, images, files, etc.)
* **Quality Evaluation**: Understanding how your application performs in real user scenarios and using evaluation criteria relevant to your product
* **Experimentation & Iteration**: Systematically versioning, testing, and deploying changes to prompts, tools, and/or models
* **Team Collaboration:** Keeping all team members on the same page when it comes to testing and understanding quality (including non-developers)
Freeplay addresses these challenges by providing a comprehensive solution for prompt and model management, observability, and evaluation that works seamlessly across modalities/data formats — including audio. And Freeplay makes it easy for both technical and non-technical team members to fully participate in the product development and optimization process.
## What You'll Be Able to Monitor
Once implemented, you'll be able to view complete user interactions in Freeplay, including:
* **Audio recordings** from the user and bot turns
* **Transcribed text** for easy review and analysis
* **LLM responses** with full context
* **Cost & latency metrics** for performance optimization
* **Evaluation results** against your quality criteria
Once implemented, you'll be able to view complete user interactions in Freeplay, including:
* Audio recordings
* Transcribed text
* LLM responses
* Cost & latency metrics
* Evaluation results
## Integration Approaches
Freeplay provides seamless integration with Pipecat to log audio interactions and LLM responses for comprehensive testing and evaluation of your voice agents.
### Option 1: Processor Integration
* How it works: Directly intercepts frames within the pipeline processing flow
* Trade-off: Adds minimal latency as processing happens inline
* Best for: Cases where you need direct frame manipulation or synchronous processing
* Documentation: [Pipecat Processors](https://pipecat-docs.readthedocs.io/en/latest/api/pipecat.processors.html)
* Full Example: [FreeplayProcessor](/developer-resources/recipes/pipecat-processor-integration)
### Option 2: Observer Integration ⭐ Recommended
* How it works: Uses callbacks to log data asynchronously in the background
* Trade-off: Zero impact on pipeline latency since logging happens outside the main flow
* Best for: Voice agents where low latency is critical
* Documentation: [Pipecat Observer Pattern](https://docs.pipecat.ai/server/utilities/observers/observer-pattern)
* Full Example: [FreeplayObserver](/developer-resources/recipes/freeplay-pipecat-observer)
We recommend the Observer pattern: Since latency is crucial for voice agents, the Observer approach ensures your audio pipeline runs at maximum speed while still capturing all necessary data for Freeplay.
## Conversation Flow in a Pipecat + Twilio + Freeplay Integration
The above image shows how Freeplay integrates into the pipeline. Regardless if you are using the Observer or Processor methods, the logic is very similar. Once both user and bot audio is received and complete, a turn can be logged to Freeplay. This flow ensures comprehensive logging while maintaining optimal performance for real-time voice interactions.
***
## Implementation Guide FreeplayObserver
### Pro Tip: AudioBufferProcessor
We highly recommend using Pipecat's [AudioBufferProcessor](https://pipecat-docs.readthedocs.io/en/latest/api/pipecat.processors.audio.audio_buffer_processor.html) alongside the FreeplayObserver. This well-tested utility:
* Formats audio data consistently for logging
* Provides reliable callbacks for conversation turn detection
* Simplifies audio handling and storage
This implementation guide uses the [Twilio-Chatbot](https://github.com/pipecat-ai/pipecat/tree/main/examples/twilio-chatbot) example provided by Pipecat—a chatbot accessed via phone application powered by Twilio. This integration allows you to log audio interactions and LLM responses to Freeplay for comprehensive testing and evaluation of your audio agents.
### Prerequisites
Before starting, make sure you have:
1. A Freeplay account set up (follow our [quick start guide](/getting-started))
2. Your prompts configured in Freeplay (follow our [prompting guide](/core-concepts/prompt-management/managing-prompts))
3. A working Pipecat + Twilio application
For a complete code example, see our FreeplayObserver Pipecat [Recipe](/developer-resources/recipes/freeplay-pipecat-observer).
### Step 1: Import Prompt Configuration from Freeplay
First, fetch your prompt configuration from Freeplay and prepare it for use in your Pipecat pipeline:
```python python theme={null}
from helpers.freeplay_frame import FreeplayProcessor # Processor example
from helpers.freeplay_observer import FreeplayObserver # Observer example
from freeplay import Freeplay, SessionInfo
# Freeplay Client
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base=os.getenv("FREEPLAY_API_BASE")
)
# Get the unformatted prompt from Freeplay
unformatted_prompt = fp_client.prompts.get(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=os.getenv("PROMPT_NAME"),
environment="latest",
)
formatted_prompt = unformatted_prompt.bind(
variables=,
history=[],
).format()
# Pass the formatted prompt to the LLM
llm = OpenAILLMService(model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema if formatted_prompt.tool_schema else None,
api_key=os.getenv("OPENAI_API_KEY"),
**formatted_prompt.prompt_info.model_parameters)
```
*Why bind the prompt? This approach allows you to fetch the prompt once and reuse it throughout the conversation, avoiding repeated API calls to Freeplay while still allowing you to add new variables and information at each conversation turn.*
### Step 2: Create Your Freeplay Observer
Create a FreeplayObserver to handle conversation memory, frame monitoring, and data logging. See the full code [here](/developer-resources/recipes/freeplay-pipecat-observer).
```python python theme={null}
# Note: see the full implementation in our Recipe section of the docs.
class FreeplayObserver(BaseObserver):
def __init__(
self,
fp_client: Freeplay,
unformatted_prompt: str = None,
template_name: str = os.getenv("PROMPT_NAME") or None,
environment: str = "latest",
variables: dict = {},
):
super().__init__()
self.start_llm_interaction = 0
self.end_llm_interaction = 0
self.llm_completion_latency = 0
self.call_id = str(uuid.uuid4()) # Has to be str to record to Freeplay
# Audio related properties
self.sample_width = 2
self.num_channels = 1
self.sample_rate = 16000
self._bot_audio = bytearray()
self._user_audio = bytearray()
self._turn_user_audio = bytearray()
self.user_speaking = False
self.bot_speaking = False
self.fp_client = fp_client
self.session = self.fp_client.sessions.create()
# Freeplay Params
self.template_name = template_name
self.environment = environment
self.unformatted_prompt = unformatted_prompt
self.variables = variables
# Conversation Params
self.conversation_id = self._new_conv_id()
self.conversation_history = []
self.most_recent_user_message = None
self.most_recent_completion = None
def _new_conv_id(self) -> str:
"""Generate a new conversation ID based on the current timestamp (this represents a customer id or similar)."""
return datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
def _reset_recent_messages(self):
"""Reset all temporary message and audio storage."""
self.most_recent_user_message = None
self.most_recent_completion = None
self._user_audio = bytearray()
self._bot_audio = bytearray()
self._turn_user_audio = bytearray()
# self.llm_completion_latency = 0
async def record_to_freeplay(self):
"""Record the current interaction to Freeplay as a new trace."""
# Create a new trace for this interaction
trace = self.session.create_trace(
input=self.most_recent_user_message,
custom_metadata={
"conversation_id": str(self.conversation_id),
},
)
# Add user message to conversation history
self.conversation_history.append(
{
"role": "user",
"content": [
{"type": "text", "text": self.most_recent_user_message},
{
"type": "input_audio",
"input_audio": {
"data": base64.b64encode(self._user_audio).decode("utf-8"),
"format": "wav",
},
},
],
},
)
# Bind the variables to the prompt
if self.unformatted_prompt:
formatted = self.unformatted_prompt.bind(
variables=self.variables,
history=self.conversation_history,
).format()
else:
# Run in executor to avoid blocking
loop = asyncio.get_running_loop()
formatted = await loop.run_in_executor(
None,
self.fp_client.prompts.get_formatted,
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=self.template_name,
environment=self.environment,
history=self.conversation_history,
variables=self.variables,
)
# Calculate latency for the LLM interaction
# Convert nanoseconds to seconds for proper timing
latency_seconds = self.get_llm_response_latency_seconds()
end = time.time()
start = end - latency_seconds
try:
print(f"_____* self._bot_audio: {len(self._bot_audio)}")
# Prepare metadata and record payload
custom_metadata = {"caller_id": self.call_id}
# Prepare assistants response message (mimicing the format of the llm provider message)
assistant_msg = {
"role": "assistant",
"content": [
{"type": "text", "text": self.most_recent_completion},
],
"audio": {
"id": self.conversation_id,
"data": base64.b64encode(self._bot_audio).decode("utf-8"),
"expires_at": 1729234747,
"transcript": self.most_recent_completion,
},
}
# Add assistant's response to conversation history
self.conversation_history.append(assistant_msg)
record = RecordPayload(
all_messages=[
*formatted.llm_prompt,
assistant_msg, # Add the assistant's response to the record call
],
session_info=SessionInfo(
self.session.session_id, custom_metadata=custom_metadata
),
inputs={},
prompt_info=formatted.prompt_info,
call_info=CallInfo.from_prompt_info(formatted.prompt_info, start, end),
trace_info=trace,
)
# Create recording in Freeplay
loop = asyncio.get_running_loop()
await loop.run_in_executor(None, self.fp_client.recordings.create, record)
# Record output to trace
await loop.run_in_executor(
None,
functools.partial(
trace.record_output,
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
output=self.most_recent_completion,
eval_results={},
),
)
print(
f"✅ Recorded interaction #{len(self.conversation_history) // 2} to Freeplay - LLM response time: {self.get_llm_response_latency_seconds():.3f}s",
flush=True,
)
# Reset only audio and current message data, keep conversation history
self._reset_recent_messages()
except Exception as e:
print(f"❌ Error recording to Freeplay: {e}", flush=True)
# Still reset audio buffers to prevent accumulation
# audio buffers are overwritten in event handler
self._reset_recent_messages()
async def make_wav_bytes(
self, pcm: bytes, sample_rate: int, voice: str, prepend_silence_secs: int = 1
) -> bytes:
"""Convert PCM audio data to WAV format with optional silence prepend."""
if prepend_silence_secs > 0:
silence_samples = int(
self.sample_rate
* self.sample_width
* self.num_channels
* prepend_silence_secs
)
silence = b"\x00" * silence_samples
pcm = silence + pcm
with io.BytesIO() as buf:
with wave.open(buf, "wb") as wf:
wf.setnchannels(self.num_channels)
wf.setsampwidth(self.sample_width)
wf.setframerate(sample_rate)
wf.writeframes(pcm)
return buf.getvalue()
async def on_push_frame(self, data: FramePushed):
src = data.source
dst = data.destination
frame = data.frame
direction = data.direction
timestamp = data.timestamp
# Create direction arrow
arrow = "→" if direction == FrameDirection.DOWNSTREAM else "←"
if isinstance(frame, LLMFullResponseStartFrame):
print(f"LLMFullResponseFrame: START {src} {arrow} {dst}", flush=True)
elif isinstance(frame, LLMFullResponseEndFrame):
print(f"LLMFullResponseFrame: END {src} {arrow} {dst}", flush=True)
elif isinstance(frame, TranscriptionFrame):
# Capture user if bot talks first
if self.most_recent_user_message is None:
self.most_recent_user_message = frame.text
elif isinstance(frame, OpenAILLMContextFrame):
messages = frame.context.messages
# Extract user message and completion from context
# NOTE: this replaces the TranscriptionFrame results, as this maps excatly what the llm recived.
user_messages = [m for m in messages if m.get("role") == "user"]
if user_messages:
self.most_recent_user_message = user_messages[-1].get("content")
completions = [m for m in messages if m.get("role") == "assistant"]
if completions:
self.most_recent_completion = completions[-1].get("content")
# Get relevant latency metrics for the LLM interaction
if (
isinstance(frame, OpenAILLMContextFrame)
and isinstance(src, LLMUserContextAggregator)
and isinstance(dst, BaseOpenAILLMService)
):
self.start_llm_interaction = timestamp
print(f"_____freeplay-observer.py OpenAILLMContextFrame START: {timestamp}")
elif isinstance(frame, LLMFullResponseEndFrame) and isinstance(
src, BaseOpenAILLMService
):
self.end_llm_interaction = timestamp
# update latency tally
self.llm_completion_latency = (
self.end_llm_interaction - self.start_llm_interaction
)
print(
f"_____freeplay-observer.py * set self.llm_completion_latency: {self.llm_completion_latency} ({self.get_llm_response_latency_seconds():.3f}s)"
)
def get_llm_response_latency_seconds(self):
"""Convert the raw nanosecond LLM response latency to seconds."""
return self.llm_completion_latency / 1_000_000_000
```
### Step 3: Configure Your Pipeline with Audio Buffering
Set up your Pipecat pipeline with the FreeplayObserver and audio buffering capabilities:
```python python theme={null}
from pipecat.pipeline.task import PipelineTask
from pipecat.pipeline.pipeline import Pipeline
from pipecat.processors.audio.audio_buffer import AudioBufferProcessor
# Create audio buffer for capturing conversation audio
audiobuffer = AudioBufferProcessor(
sample_rate=8000,
num_channels=1,
)
# Configure pipeline task with observer
task = PipelineTask(
pipeline,
params=PipelineParams(
audio_in_sample_rate=8000,
audio_out_sample_rate=8000,
allow_interruptions=True,
enable_metrics=True,
enable_usage_metrics=True,
),
observers=[freeplay_observer], # Add observer here
)
```
### Step 4: Set Up Audio Capture Callbacks & Configure Pipeline
Configure callbacks to capture audio at the optimal moments:
```python python theme={null}
from helpers.freeplay_observer import FreeplayObserver # Observer example
freeplay_observer = FreeplayObserver(
fp_client=fp_client,
unformatted_prompt=unformatted_prompt,
environment=os.getenv("FREEPLAY_ENVIRONMENT"),
)
#....Additional Pipeline Configuration...
task = PipelineTask(
pipeline,
params=PipelineParams(
audio_in_sample_rate=8000,
audio_out_sample_rate=8000,
allow_interruptions=True,
enable_metrics=True,
enable_usage_metrics=True, # This is used to track the usage of the LLM
),
observers=[
freeplay_observer
], # Use the FreeplayObserver to record the audio to Freeplay
)
# save audio bytes from user and store in freeplay_observer
@audiobuffer.event_handler("on_user_turn_audio_data")
async def on_user_turn_audio_data(buffer, audio, sample_rate, num_channels):
if audio and not freeplay_observer._bot_audio:
# aggregate user audio because this event could fire multiple times
# before bot responds
freeplay_observer._user_audio = freeplay_observer._turn_user_audio.extend(
audio
)
freeplay_observer._user_audio = await freeplay_observer.make_wav_bytes(
freeplay_observer._turn_user_audio,
sample_rate,
"user",
prepend_silence_secs=1,
)
elif audio and freeplay_observer._bot_audio:
freeplay_observer._user_audio = freeplay_observer._turn_user_audio.extend(
audio
)
freeplay_observer._user_audio = await freeplay_observer.make_wav_bytes(
freeplay_observer._turn_user_audio,
sample_rate,
"user",
prepend_silence_secs=1,
)
await freeplay_observer.record_to_freeplay()
# save audio bytes from bot and store in freeplay_observer
@audiobuffer.event_handler("on_bot_turn_audio_data")
async def on_bot_turn_audio_data(buffer, audio, sample_rate, num_channels):
# this assumes the user always speaks first and would cut off
# the first turn of the bot
if audio and not freeplay_observer._user_audio:
# aggregate bot audio because this event could fire multiple times
# before user responds
freeplay_observer._bot_audio = await freeplay_observer.make_wav_bytes(
audio, sample_rate, "bot", prepend_silence_secs=1
)
elif audio and freeplay_observer._user_audio:
freeplay_observer._bot_audio = await freeplay_observer.make_wav_bytes(
audio, sample_rate, "bot", prepend_silence_secs=1
)
await freeplay_observer.record_to_freeplay()
```
1. Add the `FreeplayObserver` to Your Pipeline
2. Initialize at Conversation Level The FreeplayObserver must be initialized at the conversation level to properly track the entire interaction flow.
3. Automatic Frame Processing As audio frames pass through the observer's on\_push\_frame method, it automatically updates the processor variables with both user and bot audio data and metadata.
4. Recording with AudioBufferProcessor Callbacks To determine the optimal timing for recording to Freeplay, we recommend using AudioBufferProcessor callbacks:
1. `on_bot_turn_audio_data` - Captures when the bot completes its audio response
2. `on_user_turn_audio_data` - Captures when the user finishes speaking
These callbacks provide the most reliable trigger points for logging complete conversation turns.
## Alternative: FreeplayProcessor Integration
### Step 1: Import Your Prompt from Freeplay & Pass to the LLMProcessor
*Note, here we get an unformatted prompt from Freeplay and then bind it, this allows us to pass it to the system and not have to make repeated calls to retrieve the llm prompt from Freeplay. The binding allows us to add new variables and information at each turn of the conversation. You can see more[here](/freeplay-sdk#get-a-prompt-template).*
```python python theme={null}
from helpers.freeplay_frame import FreeplayProcessor # Processor example
from helpers.freeplay_observer import FreeplayObserver # Observer example
from freeplay import Freeplay, SessionInfo
# Freeplay Client
fp_client = Freeplay(
freeplay_api_key=os.getenv("FREEPLAY_API_KEY"),
api_base=os.getenv("FREEPLAY_API_BASE")
)
# Get the unformatted prompt from Freeplay
unformatted_prompt = fp_client.prompts.get(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=os.getenv("PROMPT_NAME"),
environment="latest",
)
formatted_prompt = unformatted_prompt.bind(
variables=,
history=[],
).format()
# Pass the formatted prompt to the LLM
llm = OpenAILLMService(model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema if formatted_prompt.tool_schema else None,
api_key=os.getenv("OPENAI_API_KEY"),
**formatted_prompt.prompt_info.model_parameters)
```
### Step 2: Create a Freeplay Processor
The processor handles the memory of the conversation, processing of key frames, and and keeps track of information to log to Freeplay. This inherits from `FrameProcessor` in Pipecat. See the full code implementation [here](/developer-resources/recipes/pipecat-processor-integration).
```python python theme={null}
class FreeplayProcessor(FrameProcessor):
"""Logs LLM interactions and audio to Freeplay with simplified structure."""
def __init__(
self,
fp_client: Freeplay,
template_name: str,
session: SessionInfo = None,
required_information: str = None,
unformatted_prompt: PromptInfo = None,
):
super().__init__()
self.fp_client = fp_client
self.template_name = template_name
self.conversation_id = self._new_conv_id()
self.total_completion_time = 0
self.required_information = required_information
self.deepgram_latency = 0
# Audio related properties
self.sample_width = 2
self.sample_rate = 8000
self.num_channels = 1
self._user_audio = bytearray()
self._bot_audio = bytearray()
self.user_speaking = False
self.bot_speaking = False
# Freeplay related properties
self.conversation_history = []
self.session = session
self.most_recent_user_message = None
self.most_recent_completion = None
self.unformatted_prompt = unformatted_prompt
self.reset_recent_messages()
def _new_conv_id(self) -> str:
"""Generate a new conversation ID based on the current timestamp (this represents a customer id or similar)."""
return datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
def reset_recent_messages(self):
"""Reset all temporary message and audio storage."""
self.most_recent_user_message = None
self.most_recent_completion = None
self._user_audio = bytearray()
self._bot_audio = bytearray()
self.total_completion_time = 0
self.deepgram_latency = 0
async def process_frame(self, frame: Frame, direction: FrameDirection):
"""Process incoming frames and handle Freeplay logging."""
await super().process_frame(frame, direction)
# Handle LLM response frames
if isinstance(frame, (LLMFullResponseStartFrame, LLMFullResponseEndFrame)):
event = "START" if isinstance(frame, LLMFullResponseStartFrame) else "END"
print(f"LLMFullResponseFrame: {event}", flush=True)
# Handle LLM context frame - this is where we log to Freeplay
elif isinstance(frame, OpenAILLMContextFrame):
messages = frame.context.messages
# Extract user message and completion from context
user_messages = [m for m in messages if m.get("role") == "user"]
if user_messages:
self.most_recent_user_message = user_messages[-1].get("content")
completions = [m for m in messages if m.get("role") == "assistant"]
if completions:
self.most_recent_completion = completions[-1].get("content")
# Log to Freeplay when we have both user input and completion
if self.most_recent_user_message and self.most_recent_completion:
self._record_to_freeplay()
# Handle audio state changes
elif isinstance(frame, UserStartedSpeakingFrame):
self.user_speaking = True
elif isinstance(frame, UserStoppedSpeakingFrame):
self.user_speaking = False
elif isinstance(frame, BotStartedSpeakingFrame):
self.bot_speaking = True
elif isinstance(frame, BotStoppedSpeakingFrame):
self.bot_speaking = False
# # Handle audio data
elif isinstance(frame, InputAudioRawFrame):
if self.user_speaking:
self._user_audio.extend(frame.audio)
elif isinstance(frame, TTSAudioRawFrame):
if self.bot_speaking:
self._bot_audio.extend(frame.audio)
# Handle metrics for LLM completion time
elif isinstance(frame, MetricsFrame):
self.metrics = frame.data
for metric in frame.data:
if isinstance(metric, ProcessingMetricsData):
if "LLMService" in metric.processor:
self.total_completion_time = metric.value
elif isinstance(metric, TTFBMetricsData):
if "DeepgramSTTService" in metric.processor:
self.deepgram_latency += metric.value
# Pass frame to next processor
await self.push_frame(frame, direction)
```
***
Note: It is required to modify the `processes_frame` function in pipecat’s `base_llm.py` to pass along the OpenAILLMContext frame, this makes the handling easier in the FreeplayLLMLogger - `process_frame`:
```python python theme={null}
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
context = None
if isinstance(frame, OpenAILLMContextFrame):
context: OpenAILLMContext = frame.context
await self.push_frame(frame, direction) # Add this line here to pass frame along
elif isinstance(frame, LLMMessagesFrame):
context = OpenAILLMContext.from_messages(frame.messages)
elif isinstance(frame, VisionImageRawFrame):
context = OpenAILLMContext()
context.add_image_frame_message(
format=frame.format, size=frame.size, image=frame.image, text=frame.text
)
elif isinstance(frame, LLMUpdateSettingsFrame):
await self._update_settings(frame.settings)
else:
await self.push_frame(frame, direction)
....
```
***
### Step 3: Add the FreeplayProcessor To Your Pipeline
Initialize your `FreeplayProcessor` and add it as a step in your pipeline. It is recommended that you add this after the `STT` or the `audioBuffer` steps in your pipeline so that all of the information needed is available when you log to Freeplay.
```python python theme={null}
# Pass the Freeplay client to the FreeplayProcessor
freeplay_processor = FreeplayProcessor(
fp_client=fp_client,
template_name="voice-assistant",
session=session,
debug=True
)
#....Additional Pipeline Configuration...
pipeline = Pipeline(
[
transport.input(), # Websocket input from client
stt, # Speech-To-Text
context_aggregator.user(),
llm, # LLM
tts, # Text-To-Speech
freeplay_processor, # Freeplay Logger (after tts so it can capture assistant audio)
transport.output(), # Websocket output to client
audiobuffer, # Used to buffer the audio in the pipeline
context_aggregator.assistant(),
]
)
```
### Step 4: Start logging completions
Begin capturing real user interactions in Freeplay, this is a function of the FreeplayProcessor. The audio is being added to the conversation history for proper tracking.
```python python theme={null}
def _record_to_freeplay(self):
"""Record the current conversation state to Freeplay."""
# Create a new trace for this interaction
trace = self.session.create_trace(
input=self.most_recent_user_message,
custom_metadata={
"deepgram_latency": self.deepgram_latency,
},
)
self.conversation_history.append(
{
"role": "user",
"content": [
{"type": "text", "text": self.most_recent_user_message},
{
"type": "input_audio",
"input_audio": {
"data": base64.b64encode(
self._make_wav_bytes(
self._user_audio, prepend_silence_secs=1
)
).decode("utf-8"),
"format": "wav",
},
},
],
},
)
# Bind the variables to the prompt
if self.unformatted_prompt:
formatted = self.unformatted_prompt.bind(
variables={"required_information": self.required_information},
history=self.conversation_history,
).format()
else:
# Get formatted prompt. Note this adds latency to the pipeline
formatted = self.fp_client.prompts.get_formatted(
project_id=os.getenv("FREEPLAY_PROJECT_ID"),
template_name=self.template_name,
environment="latest",
history=self.conversation_history,
variables={"required_information": self.required_information},
)
# Calculate latency for the LLM interaction
start, end = time.time(), time.time() + self.total_completion_time
try:
# Prepare metadata and record payload
custom_metadata = {
"conversation_id": str(self.conversation_id),
}
# Add assistant's response to conversation history
last_message = {
"role": "assistant",
"content": [
{"type": "text", "text": self.most_recent_completion},
],
"audio": {
"id": self.conversation_id,
"data": base64.b64encode(
self._make_wav_bytes(self._bot_audio, prepend_silence_secs=1)
).decode("utf-8"),
"expires_at": 1729234747,
"transcript": self.most_recent_completion,
},
}
self.conversation_history.append(last_message)
# Create recording in Freeplay
self.fp_client.recordings.create(
RecordPayload(
project_id=os.get("FP_PROJECT_ID")
all_messages=[
*formatted.llm_prompt,
last_message, # Add the last message to the record call
],
session_info=SessionInfo(
self.session.session_id, custom_metadata=custom_metadata
),
inputs={"required_information": self.required_information},
prompt_version_info=formatted.prompt_info,
call_info=CallInfo.from_prompt_info(
formatted.prompt_info, start, end
),
trace_info=trace,
)
)
# Record output to trace
trace.record_output(
os.getenv("FREEPLAY_PROJECT_ID"),
self.most_recent_completion,
)
print(
f"Successfully recorded to Freeplay - completion time: {self.total_completion_time}s",
flush=True,
)
self.reset_recent_messages()
except Exception as e:
print(f"Error recording to Freeplay: {e}", flush=True)
self.reset_recent_messages()
```
### Helpful resources:
[Configure LiteLLM Proxy Models in Freeplay](/account-setup/configure-litellm-proxy-models-in-freeplay)
[Security Overview](/security-compliance/security-overview)
```
```
# Common Integration Patterns
Source: https://docs.freeplay.ai/practical-guides/common-integration-patterns
Implement common patterns like multi-turn chat, agent workflows, tool calls, and customer feedback.
### Multi-Turn Conversations
For chatbots and assistants, pass conversation history when fetching prompts:
```python python theme={null}
# Your conversation history
history = [
{'role': 'user', 'content': 'What is pasta?'},
{'role': 'assistant', 'content': 'Pasta is an Italian dish...'}
]
# Fetch prompt with history
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="chat-assistant",
environment="latest",
variables={"user_question": "How do I make it?"},
history=history # Freeplay handles formatting
)
```
[See full multi-turn example →](/developer-resources/recipes/continuous-chat)
### Tool Calling
```python python theme={null}
if completion.choices[0].message.tool_calls:
for tool_call in completion.choices[0].message.tool_calls:
if tool_call.function.name == "weather_of_location":
args = json.loads(tool_call.function.arguments)
temperature = get_temperature(args["location"])
tool_response_message = {
"tool_call_id": tool_call.id,
"role": "tool",
"content": str(temperature),
}
messages.append(tool_response_message)
```
```typescript typescript theme={null}
// Append the completion to list of messages
const messages = formattedPrompt.allMessages(completion.choices[0].message);
if (completion.choices[0].message.tool_calls) {
for (const toolCall of completion.choices[0].message.tool_calls) {
if (toolCall.function.name === "weather_of_location") {
const args = JSON.parse(toolCall.function.arguments);
const temperature = getTemperature(args.location);
const toolResponseMessage = {
tool_call_id: toolCall.id,
role: "tool",
content: temperature.toString(),
};
messages.push(toolResponseMessage);
}
}
}
```
[See full tool calling example →](/developer-resources/recipes/using-tools-with-openai)
### Adding Custom Metadata
Track user IDs, feature flags, or any custom data:
```python python theme={null}
# Create session with metadata
session = fp_client.sessions.create(
custom_metadata={
"user_id": "user_123",
"environment": "production",
"feature_flag": "new_ui_enabled"
}
)
```
### Logging User Feedback
Capture thumbs up/down or other user reactions:
```python python theme={null}
# After the user rates your response
fp_client.customer_feedback.update(
completion_id=completion.completion_id,
feedback={
'thumbs_up': True,
'user_comment': 'Very helpful!'
}
)
```
[See full feedback example →](/freeplay-sdk/recording-completions#log-customer-feedback)
### Tracking Multi-Step Workflows
For agents and complex workflows, use traces to group related completions:
```python python theme={null}
# Create a trace for multi-step workflow
trace_info = session.create_trace(
input="Research and write a blog post about AI",
agent_name="blog_writer",
custom_metadata={"version": "2.0"}
)
# Log each LLM call with the trace_id
# ... your LLM calls here ...
# Record final output
trace_info.record_output(
project_id=project_id,
output="[Final blog post content]"
)
```
[See full agent example →](/practical-guides/agents)
```
```
# Configure, Test & Deploy a Fallback LLM Provider
Source: https://docs.freeplay.ai/practical-guides/configuring-a-fallback-llm-provider-with-freeplay
Configure, test, and deploy fallback LLM providers to ensure reliability when your primary provider is unavailable.
API-based access to the most powerful AI models in the world is a beautiful thing, opening the door to a whole range of potentially groundbreaking applications. That said, as these LLMs become an increasingly important part of critical applications, the thought of having a single point of failure by way of a dependency on a single LLM provider is giving engineers, SREs and managers everywhere heartburn. And rightfully so, we would never tolerate that kind of brittleness in other parts of our application stack, why should LLM development be any different?
The good news is that the number of providers serving cutting edge LLMs by API is growing. This means provider diversification is possible! The bad news, prompt and model configs are not fully portable from one provider to another. Whether it be due to the RLHF process, the underlying training data or other factors, these models each have their own optimal prompting style. This means that in order to have a truly reliable fallback provider you need a prompt and model config that is continually validated against your benchmark dataset and primary provider for both latency and quality. This can be a daunting task without the right tooling and workflows in place.
Here’s how Freeplay can help you establish, maintain, and serve a fallback LLM provider. In this case we are using OpenAI as our primary provider and will configure Anthropic as a fallback provider.
## Step 1: Create a Dataset for Benchmarking
Having a labeled Dataset to test prompt, model, and pipeline changes against is critical for building a repeatable and robust LLM development process. This is also an important foundational component when configuring a fallback LLM provider
Freeplay provides in app functionality for you to label and curate dataset from real production sessions.
Alternatively, if you already have a dataset created you can upload those examples directly to Freeplay via JSONL.
## Step 2: Configure a Prompt Template and Model Config for your Fallback Provider
Freeplay’s prompt editor is an interactive playground allowing you to load in data from your datasets and compare prompt versions side by side. Here we have our primary provider prompt pulled up and as we iterate on a prompt for Anthropic's Sonnet model. We’ve loaded in a few examples from our benchmark dataset to test against.
## Step 3: Test your Fallback Provider at Scale
After we've created a fallback provider prompt template that seems to work well, we want to test it at scale and compare it to our benchmark dataset, which in this case was generated by our primary provider and human labeled. We can kick off the test run either in app from Freeplay or in code via the [Freeplay SDK](/freeplay-sdk/prompts).
It looks like our fallback provider is performing on par with our primary provider, and actually a bit better in some ways! Cost and latency are also similar so we know that we can continue meeting our external SLAs while keeping internal costs under control.
## Step 4: Deploy your Fallback Provider and Configure your Application Code Accordingly
Now that we’ve validated this new prompt and model config we need to make sure our code supports the new provider. We need to add an API key for our new provider and update our application code with our fallback strategy. Here’s a high level overview of how the fallback will work.
1. Using the prompt template from our primary provider we try making a request
2. If the request fails, we fetch the prompt template for our secondary provider from Freeplay and make a request
3. Record the results back to Freeplay
**Note that for steps 1 and 2 we could alternatively make use of Freeplay's[prompt bundling](/core-concepts/prompt-management/prompt-bundling) feature which allows us to check out prompts during our build process, such that we can read them from our local filesystem, rather than fetch them from the Freeplay server. This removes Freeplay from the critical path entirely.**
Here is what the code looks like
```python python theme={null}
# start timer for logging latency of the full chain
start = time.time()
# run semantic search
search_res, filter = vector_search(message, top_k=top_k,
cosine_threshold=cosine_threshold,
tag=tag, title=title)
# get the formatted prompt
prompt_vars = {
"question": message,
"supporting_information": str(search_res)
}
# get a formatted prompt for your primary provider
formatted_prompt = fpClient.prompts.get_formatted(
project_id=freeplay_project_id,
template_name="rag-qa",
environment="prod",
variables=prompt_vars
)
# first try making a request with your primary
try:
chat_completion = openai.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages = formatted_prompt.messages,
**formatted_prompt.prompt_info.model_parameters
)
content = chat_completion.choices[0].message.content # update messages
messages = formatted_prompt.all_messages(
{'role': chat_completion.choices[0].message.role,
'content': content}
)
except: # fetch the prompt for our fallback provider
formatted_prompt = fpClient.prompts.get_formatted(
project_id=freeplay_project_id,
template_name="rag-qa",
environment="fallback",
variables=prompt_vars
)
chat_completion = anthropicClient.messages.create(
model=formatted_prompt.prompt_info.model,
system=formatted_prompt.system_content,
messages=formatted_prompt.llm_prompt,
**formatted_prompt.prompt_info.model_parameters
)
content = chat_completion.content[0].text
messages = formatted_prompt.all_messages(
{'role': chat_completion.role,
'content': content}
)
# log latency
end = time.time()
# create an async record call payload
record_payload = RecordPayload(
project_id
all_messages=messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=start, end_time=end)
)
# record the call
completion_log = fpClient.recordings.create(record_payload)
```
## Step 5: Maintain and Update your Fallback Provider
Data naturally shifts over time. It’s important to continually test and update your fallback provider config such that it maintains parity with your primary provider overtime.
**Rest Easy!**
All your engineers, SREs and managers can now sleep easier at night knowing that in the event of a provider incident you will not just continue to serve traffic, but know you are doing so with consistent quality, cost and latency.
```
```
# Creating and Aligning Model-graded Evals
Source: https://docs.freeplay.ai/practical-guides/creating-and-aligning-model-graded-evals
Model-graded evals (aka LLM-as-a-judge) have quickly become a critical component of LLM evaluation. While far from a silver bullet, reliable model-graded evals can fill the gap between code evals and human review allowing you to scale nuanced evals with intelligence.
When using evals in a product context, we generally find people want the freedom to customize their evals, even if they start with a template. It can be appealing at first to want turnkey “model-graded evals in a box,” but as teams mature they quickly realize the need to customize their model-graded evals, or to create new evals from scratch.
Our focus at Freeplay has been giving teams the tools to see exactly what their evals are doing, customize them when needed, and improve them over time.
Read more about why custom eval suites are important: [Building an effective eval suite](https://freeplay.ai/blog/building-an-llm-eval-suite-that-actually-works-in-practice)
While customization is critical, it can be tedious. Just as prompt engineering takes a lot of iteration, model-graded eval development can take a lot of time too. That’s why we’ve made it easy to set up a repeatable process for eval iteration and optimization. In this post we will look at how to first create model-graded evals, then align them with your team’s perspective on what the right answers are, so that they mimic human preferences as closely as possible.
## An overview of Evals in Freeplay
There are 3 types of evals in Freeplay: Model-graded evals, Human labeling, and Code evals.
Model-graded evals and Human labeling are defined and managed in the Freeplay app as part of a given prompt. Today, code evals are written, managed, and executed separately in your code and then recorded to Freeplay via the [SDK](/freeplay-sdk/recording-completions#recording-code-based-evaluations) or [API](/api-reference/observability/record-completion).
Model-graded evals are executed in two scenarios:
1. **Live Monitoring**
Freeplay will sample a subset of your production traffic and automatically run your model-graded evals for you (aka “auto-evals”). Given that these have an LLM running under the hood, costs can add up, so they are not run on all your production traffic by default.
2. **Test Runs**
When executing batch tests or comparisons in Freeplay, your model-graded evals will be run 100% of the time for each example in the dataset. The expectation here is that the whole point of your testing is to quantify results, and auto-evals are key to doing this.
Now let’s look at the process of creating and aligning model-graded evals with Freeplay.
## Create and Align Model-Graded Evals in Freeplay
In this guide, we will cover the process of creating a new model-graded eval. Once you create your eval you can continue to iterate on it over time, the process of iterating on and improving alignment for an existing model-graded eval is the same.
### Step 1: Create a new Model-graded Eval
First navigate to the prompt from which you want to create an eval, and click "New evaluation".
You'll then be prompted to choose between a Human labeling criteria and a Model-graded eval. In this case we will select "Model graded".
In the next section you will give your eval a name, define the eval scale, and give the eval a description. Note, the name and description in this section are purely for human consumption. None of this information is sent to the LLM. The prompt for the LLM will be defined in the next section.
### Step 2: Define your Model-graded Eval Prompt
Now you’ll define the prompt for your model-graded eval. These are the instructions the LLM will use to score each Completion it is run on.
First you will write the evaluation prompt. Since Freeplay manages your prompt templates as well, you can easily target specific components of the underlying prompt metadata by using Mustache syntax. When running an eval for a given prompt template:
* Input variables for the Prompt template are referenced via the `{{inputs.}}` prefix
* The output is referenced via `{{output}}`
* If you’re using chat history, you can reference the history via `{{history}}`
* You can optionally do pairwise comparisons, which are useful when using ground truth datasets. Access the ground truth dataset output via `{{dataset.output}}`
This variable value is important because often times the evaluator only needs certain aspects of the target prompt. Instead of sending the entire prompt to the evaluator, you can send specific inputs that let you create nuanced evals like the following (simplified) examples:
* **Context Relevance:** Is the retrieved context from `{{inputs.context}}` relevant to the original user query from `{{inputs.question}}`?
* **Answer Similarity:** Is the new version of the prompt output from `{{output}}` similar to the ground truth value from `{{dataset.output}}` ?
* **Entailment:** Does the answer from the prompt output `{{output}}` logically entail from the provided context `{{inputs.context}}` and the prior chat history `{{history}}`?
In the screenshot below we are grading Answer Completeness, so we are targeting the input question and the output. Note, we are excluding another variable (supporting information) because it is not strictly relevant for this criteria.
Once you’ve written the base evaluation prompt, you can also:
* Choose what model you want to use for the evaluator.
* (Optionally) Define an “Evaluation rubric” so that the LLM evaluator knows exactly what each scoring label means.
* Toggle the “Enable Live Monitoring” feature on or off for the given criteria. If it’s off, your eval will only run on Test scenarios.
Once you’re done with this initial draft, you can hit “Save template” in the right hand corner and then move on to testing it with a dataset.
### Step 3: Select a dataset to test your evaluator
We will now move to testing your new evaluator and we will need a dataset to test with. The default option will use your benchmark dataset for that eval criteria.
Note: Benchmark datasets are automatically built as you label examples. Examples will be sampled from your production logs to seed the dataset. Then as you human label data, those examples will be added to this criteria-specific benchmark dataset building it up over time.
Alternatively, you can select any other pre-existing dataset that is compatible with your underlying prompt.
Click “Test and Align Eval” in the bottom corner
### Step 4: Label examples
We will now start labeling examples to determine how frequently your evaluator score aligns with your human preference. This is the crux of the alignment flow!
You will be prompted to score each example yourself. After you score an example the model-graded score will appear, as well as an explanation of the underlying reasoning.
We recommend labeling at least 10 examples but there is no minimum amount you need to label before deploying a model-graded eval.
### Step 5: Publish your model-graded eval
Once you label however many examples you want you can deploy your eval criteria by hitting deploy in the top right.
Alternatively if your alignment score is lower than you’d hope for, you can iterate on your evaluator prompt and run more alignment sessions.
### Step 6: Iterate!
Alignment is meant to be an ongoing process. As you review more data and discover more edge cases you’ll likely want to update and improve your evaluator. You can come back at any time and continue iterating on your evaluator prompt making sure it is continually aligned with your human judgement.
## Key Takeaways
Model-graded evals are most effective when they are highly tailored to your use case. But creating high-quality, customized evals takes some iteration. Freeplay aims to facilitate that process by giving you the tools to directly align your model-graded evals with your SME’s judgement.
## Helpful resources
[Curating Useful Datasets for Testing & Evaluation](/core-concepts/datasets/dataset-curation)
[Configure, Test & Deploy a Fallback LLM Provider](/practical-guides/configuring-a-fallback-llm-provider-with-freeplay)
# Multi-Turn Chatbot
Source: https://docs.freeplay.ai/practical-guides/multi-turn-chat-support
Overview of implementing multi-turn chatbots using Freeplay.
# Introduction
Many LLM applications involve more than just one-off, isolated LLM completions. For chatbots especially, they consist of multiple back-and-forth exchanges between a user and assistant. This makes chatbots unique to test and evaluate.
This document walks through how to use Freeplay to build, test, review logs, and capture feedback on multi-turn chatbots, including how to make use of a special `history` object:
* Defining `history` in prompt templates
* Managing `history` with the Freeplay SDK
* Recording and viewing chat turns in Freeplay as traces
* Managing datasets, configuring evals and automating tests that include `history`
In this document, we will refer to one back-and-forth exchange between the user and the assistant as a "turn".
## Understanding History for Chatbots
First, why does history matter when building a multi-turn chatbot?
Importantly, each exchange must be aware of all the previous exchanges in the conversation — aka the "history" — such that the LLM can give an answer that is contextually aware. Experimentation and testing with multi-turn chat must also take history into account, since any simulated test cases need to include relevant context.
Consider this series of exchanges between the user and assistant:
In Turn Two, the assistant needs to have the context from the previous turn to give a reasonable answer. "I want them to be healthier" is the user's request for healthy dinner ideas that make use of rice.
By Turn Three, the assistant needs to reference Turn 2 to know what “Give me a recipe for number 2” refers to. And so forth.
Without an understanding of the context from the conversation history, each new message would be impossible to interpret.
**Note:** While chatbots are the most common UX that uses this interaction pattern, it can apply more broadly. It can be helpful to think of **`history`** as a way to manage state or memory, since the LLM itself does not store any persistent context from one interaction to the next. Nothing restricts the use of these concepts to a chatbot UX.
# Using Freeplay with Multi-Turn Chatbots
What's different about using Freeplay with a chatbot? There are a couple important things to be aware of:
1. **Prompt Templates:** You'll define a special `history` object in a prompt template allowing you to pass conversation history at the right point.
2. **Recording Multi-Turn Sessions:** You'll record `history` with each new chatbot turn, as well as record messages at the start and end of each trace to make it easy to view the `input` and `output` (see [Traces](/freeplay-sdk/traces) documentation).
3. **Managing Datasets & Testing:** You'll curate datasets that contain `history` so you can simulate accurate conversation scenarios when testing.
4. **Configuring Auto-Evaluations:** If you're using model-graded evals, you'll be able to target `history` objects for realtime monitoring or test scenarios.
## History in Prompt Templates
### History should be configured within your Prompt Templates in Freeplay.
When configuring your Prompt Template, you will add a message of type `history` wherever your history messages should be inserted. This tells Freeplay how messages should be ordered when the prompt template is formatted.
The most common configuration would look like this:
Creating this configuration on a Freeplay prompt template would look like this:
This tells Freeplay to insert the `history` messages in between the `system` message and the most recent `user` message when formatting a prompt.
You must define `history` in a prompt template before you can pass history values at record time and have them saved properly for use in datasets, testing, etc.
**Why configure history explicitly in the prompt template?**
While it may seem redundant at first to explicitly configure the placement of `history`, it allows for the support of more varied prompting patterns. For example, you may have some predefined context that you use to seed the model each time and include multiple messages in a prompt template. In that case, a prompt template could look like the following:
This would tell Freeplay to insert `history` messages after the first Assistant/User pair, rather than directly after the System message.
## Multi-Turn Chat in Logging and Observability
Freeplay makes it easy to understand complex, multi-turn conversations in your LLM applications. Using [Traces](/freeplay-sdk#traces), you can log conversations with clear input/output pairs that mirror what your users actually see—even when multiple prompts and LLM calls are happening behind the scenes.
### Understanding Session and Trace Views
When you navigate to a Session in Freeplay's Observability dashboard, you'll see the conversation from your user's perspective. Each trace represents a single turn in the conversation—the user's question and the AI's response.
In the left sidebar, you can see the conversation structure: a Session containing multiple Traces, each with their underlying completions. The main view shows the user-facing input and output, making it easy to understand the conversation flow at a glance.
### Inspecting What's Happening Under the Hood
Click on any trace to see what's actually happening behind that single user interaction. Often, a single user-facing response involves multiple LLM calls—like retrieval, reasoning, and generation steps.
The trace detail view shows you all the completions that contributed to this response. In this example, two prompts were called to answer the user's question about prompt bundling. You can see the input, output, timing, and cost for the entire trace.
### Diving into Individual Completions
Click on any completion within a trace to examine it in detail. Here you can see the full prompt, the model's response, inputs, and the conversation history. All of this granular information gives you the most context you need to perform analysis and review the conversation.
Each completion can be added to a dataset for testing, opened in the prompt editor, or added to a review queue for human evaluation. Any customer feedback logged at the completion level automatically rolls up to the trace level, making it easy to spot which user interactions need review.
## Multi-Turn Chat Testing
### Save and modify `history` as part of datasets to simulate real conversations.
Whenever you save an observed conversation turn that includes `history`, it will be included in the dataset for future testing. You can also edit or add `history` objects to a dataset at any time in case you want to control exactly what goes into a test scenario. Auto-evals can target `history` as well for faster test analysis.
### Datasets and Test Runs
When building a chatbot, the testing unit remains at the Completion level but includes `history` when relevant. Consider this example again:
If we were to save the completion that generates Turn Two to a dataset, we would also get the preceding context from Turn One, which would exist in the `history` object for the new completion.
Subsequent Test Runs using that Test Case would treat Turn One as static, meaning it is not recomputed during the Test Run. It would be passed as context when Turn Two is regenerated so that you can simulate that exact point in the conversation when testing.
Here's a simple sample dataset row that includes several messages in the history object.
### Auto Evaluations
History can be targeted in model-graded auto-evaluation templates like any other variable using the`{{history}}` parameter. This allows you to ask questions like: *Is the current output factually accurate given the preceding context?*
```
Determine whether or not the output is factually consistent with the preceding context
The output should be deemed inaccurate if it contains any logical contradictions with
the preceding context
{{history}}
```
## Multi-Turn Chat in the SDK
When formatting your prompts you will pass the previous messages as an array to the `history` parameter. The `messages` object will have the history messages inserted in the right place in the array, as defined in your prompt template. See more details in our [SDK docs here](/freeplay-sdk#using-history-with-prompt-templates).
```python python theme={null}
previous_messages = [{"role": "user", "content": "what are some dinner ideas..."},
{"role": "assistant", "content": "here are some dinner ideas..."}]
prompt_vars = {"question": "how do I make them healthier?"}
formatted_prompt = fpClient.prompts.get_formatted(
project_id=project_id,
template_name="SamplePrompt",
environment="latest",
variables=prompt_vars,
history=previous_messages # pass the history messages here
)
print(formatted_prompt.messages)
# output:
[
{'role': 'system', 'content': 'You are a polite assistant...'},
{'role': 'user', 'content': 'what are some dinner ideas...'},
{'role': 'assistant', 'content': 'here are some dinner ideas...'},
{'role': 'user', 'content': 'how do I make them healthier?'}
]
```
```typescript typescript theme={null}
// Call LLM Function
async function callLLM(input, history, sessionInfo) {
// LLM Call occurs here....
// Record the interaction to Freeplay
await freeplay.recordings.create({
projectId,
allMessages: [...history, responseMessage], // History is passed to capture the conversation history.
inputs: { input },
sessionInfo: sessionInfo,
promptInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end)
});
return responseMessage;
}
// Multi-turn chat
// Create a new session
const session = freeplay.sessions.create({
customMetadata: { conversation_topic: "home repair" },
});
const sessionInfo = getSessionInfo(session);
// Initialize conversation history
let history = [
{
role: "user",
content: "Why isn't my sink working?",
},
];
// First interaction
console.log("User: Why isn't my sink working?");
const response1 = await callLLM("", history, sessionInfo);
// Update history with the assistant's response
history.push(response1);
// Second interaction - continue the conversation
history.push({
role: "user",
content: "Tell me more about checking the P-trap",
});
console.log("User: Tell me more about checking the P-trap");
const response2 = await callLLM("", history, sessionInfo);
// Update history again
history.push(response2);
// Third interaction
history.push({
role: "user",
content: "What tools do I need for this job?",
});
const response3 = await callLLM("", history, sessionInfo);
// In a real application, you would fetch messages from the backend
// Here we're just using the history we've built up
// New follow-up question in the restored session
history.push(response3);
history.push({
role: "user",
content: "How long should this repair take?",
});
const response4 = await callLLM("", history, sessionInfo);
}
```
You can then use that prompt and messages to make a call to your LLM provider:
```python python theme={null}
s = time.time()
chat_response = openai_client.chat.completions.create(
model=formatted_prompt.prompt_info.model,
messages=formatted_prompt.messages,
**formatted_prompt.prompt_info.model_parameters
)
e = time.time()
latest_message = chat_response.choices[0].message
```
You will then pass the full set of messages back to Freeplay on the record call:
```python python theme={null}
all_messages = [...formatted_prompt.messages, latest_message]
# record the call
payload = RecordPayload(
project_id=project_id,
all_messages=all_messages,
inputs=prompt_vars,
session_info=session,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start_time=s, end_time=e)
)
completion_info = fpClient.recordings.create(payload)
```
You can then repeat that pattern with each turn in the conversation continuing to append to and update the `all_messages` object.
An end to end code recipe can be found [here](/developer-resources/recipes/continuous-chat).
***
What’s Next
Now that you're well-versed on building multi-turn chatbots using the history object, let's learn about model and key management.
# Overview
Source: https://docs.freeplay.ai/practical-guides/overview
Step-by-step guides for implementing common AI application patterns with Freeplay.
How-To Guides provide practical, task-focused instructions for implementing specific features in your AI applications. Each guide walks you through a complete implementation with code examples and best practices.
## How These Guides Fit Together
* **[Core Concepts](/core-concepts/observability/observability-dashboard)**: Understand *what* Freeplay features do and why they matter
* **How-To Guides** (this section): Learn *how* to implement specific patterns step-by-step
* **[Developer Resources](/developer-resources/overview)**: Reference documentation for SDKs, APIs, and integrations
* **[Code Recipes](/developer-resources/recipes/overview)**: Complete, runnable code examples you can copy and adapt
## Available Guides
### Integration Patterns
Quick reference for multi-turn chat, tool calling, metadata, and feedback
Structure agent workflows with traces and multi-step reasoning
### Chat & Conversations
Maintain conversation context across multiple exchanges
Implement function calling with OpenAI, Anthropic, and other providers
### Advanced Topics
Work with images, audio, and other non-text content
Create and calibrate model-graded evaluations
Handle streaming LLM responses in your application
Configure backup providers for resilience
### Voice & Specialized
Create voice-enabled AI applications using Pipecat and Twilio
## Next Steps
Once you've implemented a pattern, explore [Code Recipes](/developer-resources/recipes/overview) for complete examples you can run and adapt, or dive into the [SDK documentation](/freeplay-sdk/setup) for detailed API reference.
# Tool Calls
Source: https://docs.freeplay.ai/practical-guides/tools
Managing tool schema, recording tool calls with Freeplay
# Overview
Tools allows LLMs to call external services. A tool schema describes the tool's capabilities and parameters. When invoked, the LLM provider responds with a tool call with the specified parameters from the schema. You can learn more about OpenAI tools [here](https://platform.openai.com/docs/guides/function-calling) and Anthropic tools [here](https://docs.anthropic.com/en/docs/build-with-claude/tool-use).
## How does Freeplay help with tools?
Freeplay supports the complete lifecycle of working with tools - from managing tool schemas and recording tool calls to surfacing detailed tool call information in observability and testing. This comprehensive approach enables rapid iteration and testing of your tools.
With the Freeplay web app, you can define tool schemas alongside your prompt templates. The Freeplay SDK formats the tool schema based on your LLM provider. The SDK also supports recording tool calls, associated schemas, and responses from tool calls. You have complete control over how much of this functionality you want to use.
### Managing your tool schema with Freeplay
Freeplay enables you to define tools in a normalized format alongside your prompt template. Freeplay will then handle the translation of your tool definitions across providers so you can smoothly navigate across different providers. Simply provide a Name, Description and a JSON Schema to represent the parameters.
For example here is how you would define the parameters for a tool that fetches the weather
```json json theme={null}
{
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state e.g. San Francisco, CA"
},
"unit": {
"type": "string",
"enum": [
"c",
"f"
]
}
},
"additionalProperties": false,
"required": [
"location",
"unit"
]
}
```
*properties* represent the parameters of the tool and will be reflected in the arguments of the resulting tool call.
Each parameter has a *type* and either a *description* or an *enum* representing the possible options.
If your tool does not require any parameters to be passed you will want to set an empty schema like this
```json json theme={null}
{
"type": "object",
"properties": {}
}
```
**Here is how you add a Tool to a Prompt Template**
* When adding or editing a prompt template with a supported LLM provider, you will see an "Manage tools" button.
*
Enter a name, description, and parameters for your tool
* Click on "Add tool" to add the tool to the prompt template. The prompt template will be in draft mode for you to run interactively in the editor. From there you can click save and create a new prompt template version with your tool schema attached.
### Using the tool schema and recording tool calls
You can use the Freeplay SDK to fetch the tool schema as part of prompt retrieval and automatically format it for your LLM provider.
```python python theme={null}
# Fetch prompt template and tool schema from Freeplay
template_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name="your-prompt",
environment="latest",
variables={"user_input": "Hi"}
)
# Pass prompt, model, tools, and parameters to OpenAI client
completion = openai_client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
```
### Logging Tool Calls to Freeplay
When building agents that use tools, tool calls are always recorded as the output of an LLM call. This works by default as long as you properly record the output messages of the LLM call.
You can also add explicit tool spans to provide more data about tool execution, including latency and other metadata. These are recorded as a Trace with `kind='tool'` and linked to the parent completion.
#### Default: Tool calls in completions
Tool calls are recorded as the output from the LLM call, just as you would any other [completion](/core-concepts/observability/sessions-traces-and-completions#completions). When you call `formatted_prompt.all_messages()`, the LLM's tool call output is concatenated into the message history. After executing the tool, you add the tool result as a subsequent message, which becomes an input to the next LLM call.
```python python theme={null}
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name='my-openai-prompt',
environment='latest',
variables=input_variables
)
start = time.time()
completion = client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# Append the completion to list of messages even if it is a tool call message
messages = formatted_prompt.all_messages(completion.choices[0].message)
session = fp_client.sessions.create()
fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
tool_schema=formatted_prompt.tool_schema
)
)
```
This will create a view that looks like this
The result of the tool call is then added to the message history and passed as input to the next LLM call.
#### Adding explicit tool spans
You can add explicit tool spans to provide more data about tool execution. This is useful for:
* Debugging complex agent workflows with many tool calls
* Measuring tool execution timing separately from LLM latency
* Surfacing tool behavior more prominently in observability dashboards
Tool spans are logged **in addition to** the tool calls appearing in the message history—they provide extra visibility, not a replacement.
To link tool calls to the completion that generated them, use the
`completion_id` returned from the `recordings.create()` method as the
`parent_id` when creating the tool span.
```python python theme={null}
formatted_prompt = fp_client.prompts.get_formatted(
project_id=project_id,
template_name='my-openai-prompt',
environment='latest',
variables=input_variables
)
start = time.time()
completion = openai_client.chat.completions.create(
messages=formatted_prompt.llm_prompt,
model=formatted_prompt.prompt_info.model,
tools=formatted_prompt.tool_schema,
**formatted_prompt.prompt_info.model_parameters
)
end = time.time()
# Append the completion to list of messages even if it is a tool call message
messages = formatted_prompt.all_messages(completion.choices[0].message)
session = fp_client.sessions.create()
# Record the LLM completion and get the completion_id
record_response = fp_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=messages,
session_info=session.session_info,
inputs=input_variables,
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info, start, end),
tool_schema=formatted_prompt.tool_schema
)
)
completion_id = record_response.completion_id
# Create tool spans as children of the completion
if completion.choices[0].message.tool_calls:
for tool_call in completion.choices[0].message.tool_calls:
name = tool_call.function.name
args = json.loads(tool_call.function.arguments)
# Link this tool span to the completion that triggered it
tool_trace = session.create_trace(
input=args,
name=name,
kind='tool',
parent_id=uuid.UUID(completion_id)
)
# Execute and record the tool result
tool_result = tool_handler(name, args)
tool_trace.record_output(project_id=project_id, output=tool_result)
```
```typescript typescript theme={null}
const formattedPrompt = await fpClient.prompts.getFormatted({
projectId,
templateName: "my-openai-prompt",
environment: "latest",
variables: inputVariables
});
const startTime = new Date();
const completion = await openaiClient.chat.completions.create({
messages: formattedPrompt.llmPrompt,
model: formattedPrompt.promptInfo.model,
tools: formattedPrompt.toolSchema,
...formattedPrompt.promptInfo.modelParameters
});
const endTime = new Date();
// Append the completion to list of messages even if it is a tool call message
const allMessages = formattedPrompt.allMessages(completion.choices[0].message);
const session = fpClient.sessions.create();
// Record the LLM completion and get the completion_id
const recordResponse = await fpClient.recordings.create({
projectId,
allMessages,
sessionInfo: getSessionInfo(session),
promptVersionInfo: formattedPrompt.promptInfo,
callInfo: {
provider: formattedPrompt.promptInfo.provider,
model: formattedPrompt.promptInfo.model,
startTime,
endTime,
modelParameters: formattedPrompt.promptInfo.modelParameters
},
toolSchema: formattedPrompt.toolSchema
});
const completionId = recordResponse.completionId;
// Create tool spans as children of the completion
const toolCalls = completion.choices[0].message.tool_calls;
if (toolCalls) {
for (const toolCall of toolCalls) {
if (toolCall.type !== "function") continue;
const name = toolCall.function.name;
const args = JSON.parse(toolCall.function.arguments);
// Link this tool span to the completion that triggered it
const toolTrace = session.createTrace({
input: args,
name,
kind: "tool",
parentId: completionId
});
// Execute and record the tool result
const toolResult = await toolHandler(name, args);
await toolTrace.recordOutput(projectId, toolResult);
}
}
```
This creates a proper hierarchy in the trace view. Here's an example of a
completion that triggered three tool calls:
```
Session
└── Trace (user request → final response)
└── Completion (LLM call that requested tools)
├── Tool Span: search_web (query → results)
├── Tool Span: read_file (path → contents)
└── Tool Span: execute_code (code → output)
```
### Using tools in a test run
Check out [this](/developer-resources/recipes/run-a-test-with-tools-programmatically) recipe that demonstrates how to use tools programmatically in a test run.
## Code recipes
For complete code examples of tool calling with different providers:
* [Using Tools with OpenAI](/developer-resources/recipes/using-tools-with-openai)
* [Using Tools with Anthropic](/developer-resources/recipes/using-tools-with-anthropic)
# Multimodal Data
Source: https://docs.freeplay.ai/practical-guides/working-with-multi-modal-data-in-freeplay
Freeplay supports multimodal data in your prompts and completions, allowing you to work with images, audio files, and documents alongside text. This guide explains how to use multimodal features throughout the Freeplay platform.
## Overview
Many LLM applications now leverage multi-modal data types beyond just text. For mutimodal models, product images, charts, PDFs, and even audio can provide critical context to generate better responses.
**Quick start**
To get started using multimodal support within Freeplay follow these steps:
1. Define media variables in prompt templates (similar to other Mustache variable)
2. Upload and test sample files in the Freeplay prompt editor
3. Update your code to handle media files with the Freeplay SDK
4. Record and view multimodal interactions in Freeplay
5. Save recorded examples to your data datasets
This document walks through more details of how to use Freeplay to build, test, review logs, and capture feedback for your multimodal LLM applications, including how to make use of media inputs.
## Introduction
### Understanding Multimodal Data for LLMs
Multimodal models can process and analyze different types of data such as images, audio, and documents alongside text. This allows your LLM applications to "see," "hear," and "read" just like humans do. Consider these examples of how multimodal data enhances LLM applications:
**Image + Text**
* User uploads a product image with a defect and asks: "What's wrong with my product?"
* The LLM can see the image, identify the issue, and provide a relevant response.
**Document + Text**
* User uploads a financial report and asks: "Summarize the key findings in this report."
* The LLM can analyze the document contents and generate an accurate summary.
**Audio + Transcript**
* User uploads a phone call recording and asks: "Describe the tone of this call and summarize the key points".
* The LLM can analyze the audio and provide tonal analysis and generate a more accurate summary with that in mind.
Using multimodal inputs allows the LLM to interpret the additional context which can help improve your LLM system outputs, both of the above examples are able to provide much more detailed responses due to using multimodal.
## Using Freeplay with Multimodal Data
What's different about using Freeplay with multimodal data? There are a couple important things to be aware of:
1. **Prompt Templates**: You'll define media variables in a prompt template allowing you to pass image, audio, or document data at the right point. This can only be done with models that support multi-media inputs.
2. **Recording Multimodal Data**: You'll record media inputs with each completion, making it possible to view the original inputs alongside the LLM's responses during review.
3. **Media in History**: You can record media as part of history, helping you preserve key context and inputs passed within your system.
## Media Variables in Prompt Templates
### **Media variables should be configured within your Prompt Templates in Freeplay.**
When configuring your Prompt Template, you will add media variables to user or assistant messages. This tells Freeplay where to insert image, audio, or document data when the prompt template is formatted.
The most common configuration would look like this:
1. When editing or creating a prompt template in the playground, click the "Add media" button next to the prompt section type
2. Note: Media can only be added to user or assistant message types
3. Enter a variable name for your media input (e.g., `product_image`, `support_document`)
4. Select the media type (file, image or audio, types depend on the models support)
This tells Freeplay to insert the media input at that specific location in the message when formatting a prompt.
You must define media variables in a prompt template before you can pass media inputs at record time and have them saved properly for use in datasets.
## Multimodal Data in Freeplay
### Observability and Completions
### **View original media inputs alongside LLM responses in Freeplay's interface.**
When reviewing completions in Freeplay, you'll be able to see the original images, documents, or audio files that were included in the prompt. This provides essential context when evaluating model performance.
In the Freeplay observability interface, completions that include media inputs will display the media alongside the text inputs and outputs. This makes it easy to understand the full context of each interaction.
When clicking into a specific completion, you can see:
* The full prompt including all media inputs
* The model's response
* Evaluation scores and feedback
This visibility is crucial for understanding how your multimodal LLM is performing and identifying areas for improvement.
### Testing Multimodal Prompts with Real Data
Freeplay enables rapid prompt iteration by allowing you to load completions with multimodal data directly into the prompt playground. This streamlined workflow lets you test new prompt versions against real production data without leaving the editor interface.
Beyond production completions, you can also pull data from existing datasets or upload new examples for testing. This comprehensive testing approach—combining production data, curated datasets, and custom examples—accelerates your prompt development cycle and helps you deliver better AI experiences to your customers.
## Testing Workflows with Multimodal
## Multimodal Data in the SDK
### **Create a`media_inputs` map when formatting prompts via the SDK.**
When using the Freeplay SDK, you'll create a map of media variable names to their corresponding data, then pass this map to the `get_formatted` method.
### Creating Media Inputs
### Using the Media Input Map
To create the media map, import the proper type from `freeplay.resources.prompts` Then create a map of the variable name in your Freeplay prompt template to the data associated with it. In the examples below, the variable names are `product_image`, `legal_document`, and `voice_recording` .
Freeplay accepts media in one of two formats: base64-encoded data or via URL. Depending on which format you choose, you'll need to adjust how you create the media input map accordingly. See the examples below for implementation details.
To work with multimodal data in your code, follow these steps:
1. Create a media input map (either `MediaContentUrl` or `MediaContentBase64`)
2. Pass it to the `get_formatted` method
3. Include it when recording the completion
Here's how to create a media input map:
```python python theme={null}
# New imports
# Note this is in version 5.2.0 and up of the api, for older versions use freeplay.resources.prompts
from freeplay.model import MediaInputBase64, MediaInputMap, TextBlock
# Create media input map for an image
media_inputs = {
'product_image': MediaInputBase64(
type="base64",
content_type="image/jpeg",
data=encode_image_data("product.jpg") # Your function to encode image
)
}
# For a PDF document
media_inputs = {
'legal_document': MediaInputBase64(
type="base64",
content_type="application/pdf",
data=encode_file_data("contract.pdf") # Your function to encode PDF
)
}
# For audio
media_inputs = {
'voice_recording': MediaInputBase64(
type="base64",
content_type="audio/mpeg", # change audio types here
data=encode_audio_data("recording.mp3") # Your function to encode audio
)
}
###########################################
# Using URLs
###########################################
media_inputs = {
'product_image': MediaContentUrl(
type="base64",
content_type="image/jpeg",
url="https://localhost/product.jpeg" # Your function to encode image
)
}
# For a PDF document
media_inputs = {
'legal_document': MediaContentUrl(
type="base64",
content_type="application/pdf",
url="https://localhost/contract.pdf" # Link to pdf file
)
}
# For audio
media_inputs = {
'voice_recording': MediaContentUrl(
type="base64",
content_type="audio/mpeg", # change audio types here
url="https://localhost/audio.mpeg" # link to audio file
)
}
```
```typescript typescript theme={null}
import Freeplay, {
getCallInfo,
getSessionInfo,
MediaInputMap,
MediaContentUrl,
SessionInfo,
} from ...
// Image data
const media: MediaInputMap = {
"image-one": {
type: "url",
url: "https://localhost/image",
},
"image-two": {
type: "base64",
content_type: "image/jpeg",
data: "some-base64-data",
},
};
// Audio Data
const media: MediaInputMap = {
"voice_recording_1": {
type: "base64",
content_type: "audio/mpeg",
data: audioData,
},
"voice_recording_2": {
type: "url",
content_type: "audio/mpeg",
url: "https://localhost/audio.mpeg",
},
};
// File
const media: MediaInputMap = {
"legal_document_1": {
type: "base64",
content_type: "application/pdf",
data: documentData,
},
"legal_document_2": {
type: "url",
content_type: "application/pdf",
url: "https://localhost/contract.pdf",
},
};
```
### Getting Formatted Prompt with Media
When calling the Freeplay API to get a formatted prompt, include your media inputs:
```python python theme={null}
formatted_prompt = freeplay_client.prompts.get_formatted(
project_id=project_id,
template_name="multimodal-prompt",
environment="latest",
variables=input_variables,
media_inputs=media_inputs # Include your media inputs here
)
```
```typescript typescript theme={null}
const formattedPrompt =
await freeplay.prompts.getFormatted({
projectId,
templateName,
environment: "latest",
variables: input_variables,
media, // Pass in the media map to the formatted prompt
});
```
### Recording Completions with Media
When recording the completion, make sure to include the media inputs:
```python python theme={null}
record_response = freeplay_client.recordings.create(
RecordPayload(
project_id=project_id,
all_messages=[
*formatted_prompt.llm_prompt,
{"role": "assistant", "content": response_content}
],
session_info=session_info,
inputs=input_variables,
media_inputs=media_inputs, # Include your media inputs here
prompt_version_info=formatted_prompt.prompt_info,
call_info=CallInfo.from_prompt_info(formatted_prompt.prompt_info,
start_time, end_time),
)
)
```
```typescript typescript theme={null}
await freeplay.recordings.create({
projectId,
allMessages: [
...(formattedPrompt.llmPrompt || []),
{
role: "assistant",
content,
},
],
inputs: input_variables,
mediaInputs: media, // Pass the media here
sessionInfo: session_info,
promptInfo: formattedPrompt.promptInfo,
callInfo: getCallInfo(formattedPrompt.promptInfo, start, end)
});
```
## Media Support
### Supported Media Types
Freeplay supports the following media types:
* Images - JPG,JPEG,PNG, WebP
* Audio - WAV, MP3
* Documents - PDFs
Note: Support for specific file types depends on the model provider's capabilities. Please reach out to [privacy@freeplay.ai](mailto:privacy@freeplay.ai) if you’re interested to use other data types.
### Supported Sizes
We support a total request size of up to 30 mb. If your file/data is over that limit it will not work within the Freeplay Application.
### Supported Providers
Multimodal functionality is supported today by default with:
* [OpenAI](https://platform.openai.com/docs/models)
* [Claude (Anthropic)](https://www.anthropic.com/api)
* [Gemini (Google)](https://cloud.google.com/use-cases/multimodal-ai?hl=en)
Please reach out to [support@freeplay.ai](mailto:support@freeplay.ai) if you’re interested in using other models.
## Best Practices
* **Keep file sizes reasonable**: While Freeplay supports various file sizes, providers may have limits on the size of media files they can process, this can also drive up costs.
* **Test & Monitor thoroughly**: Multimodal models may perform differently with various types of images, audio quality, or document formats, Freeplay allows for rapid testing, review and iteration to ensure your product performs as expected.
* **Combine media types**: For complex use cases, you can include multiple media inputs of different types in the same prompt such as documents and images.
* **Iterate regularly**: Regularly review completions with media inputs to ensure your model is interpreting the media correctly.
Now that you're well-versed on working with multimodal data in Freeplay, you can enhance your LLM applications with rich, contextual understanding of various media types.
# Changelog
Source: https://docs.freeplay.ai/resources/changelog
What's new in Freeplay: Platform updates, SDK releases, and API changes.
Stay up to date with the latest improvements to Freeplay. This changelog covers platform features, SDK releases, API additions, and bug fixes that improve what you can do with Freeplay.
**For developers**: Watch for SDK and API updates that may require code changes. Breaking changes are clearly marked.
## February 2026
Integrate Freeplay capabilities into MCP-compatible tools and workflows with our experimental Model Context Protocol server, now available as a public repository.
[View on GitHub →](https://github.com/freeplayai/freeplay-mcp-server)
Additional links:
* [https://github.com/freeplayai/freeplay-plugin](https://github.com/freeplayai/freeplay-plugin)
* [https://github.com/freeplayai/freeplay-skills](https://github.com/freeplayai/freeplay-skills)
### **🏚️**Project Home Page
We’ve added a new Home page to every project with key metrics, insights about the project, and bookmark-able metrics. It’s a much faster way to understand what’s happening in your project. See this [Loom](https://www.loom.com/share/f4d77899e6924ad5b5410d8e544f35ee) for more information.
### 🤖 Models
**Claude Opus 4.6** — Added the newest Claude Opus 4.6 to Freeplay's prompt playground.
**Claude Haiku 4 media support** — Full image and file upload support for Anthropic Claude Haiku 4 models via both direct Anthropic API and AWS Bedrock.
### 🔧 API
**User management endpoints** — Filter deleted users via `include_deleted` query parameter and reactivate soft-deleted users through new admin endpoints.
**Insights Endpoints** - You can now get insights from Freeplay by using the `/project/{project_id}/insights` api endpoint.
**Insight filtering in search** — Search API supports filtering by `insight_id` across review themes and evaluation insights.
### 📚 Documentation
**Filtering and search documentation** — New documentation explaining tokenization behavior, phrase matching, field-type specific search, and 'contains' semantics in the observability UI.
### 🐛 Bug fixes / Improvements
* UI improvements including scrollable evaluation explanations, better test run comparison alignment, and standardized tab styling.
## January 2026
Query your observability data programmatically with three new search endpoints for sessions, traces, and completions. Build complex queries with compound filters (AND, OR, NOT), paginate through results, and use advanced filtering by eval score, cost, latency, metadata, and more.
[View Search API Operators →](/openapi/search-api-operators)
Major refresh to our documentation including an OpenAPI spec, a new `llms.txt` as the starting place for coding agents, and restructured SDK documentation. We've also added this changelog. Let us know what you think.
[Explore the docs →](/getting-started/freeplay-introduction)
### 📦 SDK
**Google GenAI tool schema update** — Define tool schemas using `GenaiFunction` and `GenaiTool` dataclasses in Python, with full TypeScript type safety in Node.
**Python SDK v0.5.5–0.5.6** — Standardized documentation, improved variable naming conventions, and reorganized capabilities. (See full [Python SDK changelog](https://github.com/freeplayai/freeplay-python/blob/main/CHANGELOG.md))
**Node SDK v0.5.2–0.5.3** — Revamped README for open source release with improved examples and documentation. (See full [Node SDK changelog](https://github.com/freeplayai/freeplay-node/blob/main/CHANGELOG.md))
### 🐛 Bug fixes
* Fixed tool call import when saving test cases from completions
Python and Node.js SDKs are now available under the Apache-2.0 license.
* [Python SDK](https://github.com/freeplayai/freeplay-python)
* [Node SDK](https://github.com/freeplayai/freeplay-node)
### 🖥️ Platform
**Run all evaluations button** — Trigger evaluation runs for all completions and traces in a session with a single click.
**CSV export for traces** — Export trace data directly from the observability view for offline analysis.
**Bulk dataset operations** — Select multiple rows in datasets to bulk delete, duplicate, or move test cases. Sort by name, compatibility, or creation date with shareable URL parameters.
### 🔧 API
**Model Management API** — Programmatically create, read, update, and delete model configurations through new CRUD endpoints.
**OpenAPI specification** — Complete schema with descriptions for all 67 API endpoints, accessible in the Freeplay app with interactive playground. [View API Reference →](/developer-resources/api-reference)
### 📦 SDK
**Metadata updates** — Update session and trace metadata after creation via `client.metadata.updateSession()` and `client.metadata.updateTrace()` in Python, Node, and JVM SDKs.
## December 2025
Our new AI agent works alongside your human reviewers to perform real-time root cause analysis, automatically surfacing patterns and actionable improvements as reviews happen.
[Learn more →](https://freeplay.ai/blog/introducing-review-insights-turn-human-notes-labels-into-action-with-freeplay-s-agent)
Define custom searches, then automatically run evaluations, add results to review queues or datasets, or trigger Slack notifications. Build weekly review queues of low-scoring logs, curate important results, or get alerts for evaluation failures.
[See the guide →](/core-concepts/observability/automations)
### 🖥️ Platform
**Updated session view** — Session cards now display evaluation scores, notes, auto-categorization results, and multiselect values. Tree view includes colored performance icons (green → red) to quickly identify problem areas. [View documentation →](/core-concepts/observability/sessions-traces-and-completions)
### 🤖 Models
**New models** — GPT-5.2, Gemini Pro 3 Flash Preview, Gemini 3 (with `thinking_level` parameter), and Mistral 3 series.
**LiteLLM for evaluations** — LiteLLM models now supported for automated evaluations.
### 🏢 Enterprise
**Directory sync** — Automatically sync users and groups from your identity provider via SCIM. Map directory groups to Freeplay roles with automatic provisioning and deprovisioning. [Learn more →](/account-setup/sso-and-scim#single-sign-on-sso-via-saml)
### 🐛 Bug fixes
* Fixed Bedrock provider `tool_result` handling
* Fixed CSV export timeout issues
* Improved text search with exact phrase matching
### 🖥️ Platform
**Create evaluations from review themes** — When you find a common issue, turn it into an LLM judge evalution directly from review themes so you can catch the issue next time it happens.
**Prompt optimization from review themes** — Use learning from a review to launch a targeted AI-powered prompt optimization experiment, using reviewed sessions as a data source.
**Slack integration** — Connect Slack workspaces to receive automation notifications with direct links to filtered views.
### 🤖 Models
**New models** — Claude Opus 4.5 and GPT-5.1 available in playground and for automated evaluations.
### 🐛 Bug fixes
* Fixed Anthropic Bedrock tool call handling with tool call history
## November 2025
Native support for LangGraph workflows, Vercel AI SDK, and Google Agent Development Kit with full observability and prompt management.
[View integrations →](/developer-resources/overview)
### 🖥️ Platform
**Tool span tracing** — Log tool calls as explicit spans with `kind="tool"`. Add custom names for clearer identification in traces. [See the Tools guide →](/practical-guides/tools)
**Review Agent (Beta)** — Automatically surfaces review themes by analyzing patterns across your review queues. Includes auto-assignment, automatic status updates, and keyboard shortcuts.
**One-click curation** — Add completions to review queues or datasets directly from session view. Edit inputs/outputs and create golden test cases in one step.
**Multimodal dataset history** — Create test cases with images and media across multiple conversation turns.
### 📦 SDK
**Node.js/TypeScript SDK v0.5.2** — Official release with full support for prompts, sessions, traces, recordings, and test runs.
```bash theme={null}
npm install freeplay
```
**Python SDK v0.5.4** — Improved package management, documentation, and multimodal data handling.
```bash theme={null}
pip install freeplay
```
[Get started →](/freeplay-sdk/setup)
## October 2025
End-to-end structured output support across Python, Node.js, and JVM SDKs. Define output schemas in prompt templates for validated JSON responses with OpenAI and Azure providers.
[Learn more →](/core-concepts/prompt-management/structured-outputs/structured-outputs)
### 🖥️ Platform
**Review queues for traces** — Systematically evaluate traces with customizable themes and automatic categorization. Trigger evaluations from OpenTelemetry data streams. [Learn more →](/core-concepts/review-queues)
### 🔧 API
**Prompt Templates API** — Create, read, update, and delete prompt versions programmatically. Update environment assignments through SDK methods. [View API Reference →](/developer-resources/api-reference)
**Environments API** — Full CRUD operations for deployment environments. [Learn more →](/core-concepts/prompt-management/deployment-environments)
### 🤖 Models
**New models** — Claude Haiku 4.5, Nova Models on AWS Bedrock (with multimedia and tool calls), and Gemini updates with fixed tool use.
**AWS Bedrock Converse API** — Comprehensive support including tool calling and multimedia inputs. [See the recipe →](/developer-resources/recipes/call-anthropic-on-bedrock)
### 🐛 Bug fixes
* Fixed sessions not displaying in review queue context
* Fixed observability date filter functionality
* Fixed duplicate test case updates
* Fixed span indentation for childless spans
* Fixed Anthropic cost calculation with OpenInference
### 🖥️ Platform
**Dataset curation improvements** — Edit outputs when saving logs to datasets for better ground truth. View ground truth in playground after loading datasets.
**Bulk auto-evaluations** — Run evaluations across multiple completions at once. Auto-trigger when completions are added to review queues.
**Trace display options** — Toggle between plain text, Markdown, and JSON formats for inputs and outputs.
### 🔧 API
**Dataset APIs** — Endpoints for getting, updating, and deleting prompt and agent datasets. OpenAPI docs support live testing in browser. [Explore →](/developer-resources/api-reference)
### 🔧 API
**Dataset Management APIs** — POST endpoints for creating datasets with configurable input names, media inputs, and history support.
### 🖥️ Platform
**OpenTelemetry expansion** — Capture Freeplay-specific attributes including provider/model info, environment tags, prompt/test IDs, metadata, and tool schemas. [Learn more →](/developer-resources/integrations/tracing-with-otel)
### 🐛 Bug fixes
* Fixed agent cost calculation showing \$0.00 for top-level costs
* Fixed auto-evaluations not working on traces
* Fixed auto-evaluation failures for criteria without `eval_prompt`
### 🔧 API
**Delete API for prompt template versions** — Programmatic removal through the v2 API.
### 🐛 Bug fixes
* Fixed "mark as best" auto-navigation behavior
* Restored next/previous navigation on filtered test runs
## September 2025
Automatically categorize logs using your own classification criteria—similar to LLM judges but for content analysis. Identify issue types that lead to evaluation failures or negative feedback.
[Learn more →](/core-concepts/evaluations/auto-categorization)
AI-powered optimization uses your live logs, evaluations, human labels, and customer feedback to recommend better prompts—and can update prompts for new models.
### 🐛 Bug fixes
* Fixed Gemini tool call correlation with OpenInference instrumentation
* Fixed next/previous navigation on filtered test runs
* Fixed test run execution with Gemini models
* Improved error messages for malformed OpenTelemetry data
### 🖥️ Platform
**Multi-modal template variables** — Access all variables from multi-modal prompts when creating datasets or configuring evaluations.
### 🖥️ Platform
**Selective evaluation control** — Choose which evaluations run during tests via UI or SDK for targeted testing and cost savings.
**Test run comparison** — Clearer cost and latency metrics rolled up at prompt and trace levels.
**Multimodal evaluations** — Target image and audio attachments with auto-evaluators. Models automatically filtered by supported media types.
**Project-level data retention** — Set shorter retention windows for sensitive projects. [Learn more →](/security-compliance/data-retention-policy)
## August 2025
**SDK breaking changes** — These changes enable optional prompt management, OTel logging support, nested traces, and multi-modal dataset management.
1. `project_id` is now the first required argument to `RecordPayload`:
```python theme={null}
RecordPayload(project_id=project_id, ...)
```
2. `PromptInfo` renamed to `PromptVersionInfo` (now optional):
```python theme={null}
RecordPayload(
project_id=project_id,
prompt_version_info=formatted_prompt.prompt_info,
...
)
```
### 🖥️ Platform
**Media input support** — Create and upload media-backed test cases with automatic type inference.
**Tree-based session interface** — Left-hand tree navigation, resizable review panel, and deep-linking for shareable session URLs.
**Multi-project service accounts** — Service accounts can now access multiple projects.
### 🤖 Models
**Tool calling expansion** — Vertex AI and Gemini tool calling, including native support in JVM SDK.
### 🐛 Bug fixes
* Fixed navigation stale selections during pagination
* Fixed Gemini test runs with proper message type conversion
* Fixed table flickering and media preview reloading
* Improved error handling for API keys from deleted users
### 🤖 Models
**New models** — GPT-5 available in playground and for evaluations. Claude Opus 4.1 and GPT-OSS models (20B/120B) can be added for your preferred inference provider.
Create trace-level LLM judges in the Freeplay UI to evaluate full agent behavior. Filter and graph agent evals separately from prompt-level evals.
[Learn more →](/practical-guides/agents)
### 🖥️ Platform
**Playground diff view** — Row-level change comparison for any two columns to compare prompt iterations.
**Prompt optimization (experimental)** — Use log examples, eval scores, human labels, and feedback to suggest prompt improvements.
**Test results filtering** — Filter graphs and test case rows together to explore metrics for different data slices.
### 🐛 Bug fixes
* Fixed filtering operators to respect numeric types (greater than, less than) instead of only string operators
## July 2025
### 🔒 Security
**WorkOS authentication** — Upgraded authentication for enhanced security and smoother logins.
### 🔧 API
**User-scoped API keys** — Full API use with private projects:
* **Private projects** → Accessible only to API keys from project members
* **Public projects** → Accessible to all API keys
## June 2025
Define and run agent-level evaluations, curate datasets for agent testing, compare agent versions, and simplified trace observability.
Systematically review and annotate AI outputs with customizable workflows.
Search across all LLM logs with instant results and trend visualizations.
Turnkey private hosting in any cloud for enterprise data residency requirements.
# Glossary
Source: https://docs.freeplay.ai/resources/glossary
Definitions of key terms and concepts used throughout Freeplay documentation.
This glossary provides definitions for the key terms and concepts you'll encounter when using Freeplay. Terms are organized by category to help you understand how they relate to each other.
A top-level workspace in Freeplay that contains all your prompt templates, sessions, datasets, evaluations, and configurations. Projects are identified by a unique `project_id` and represent a distinct AI application or use case. All other entities in Freeplay (sessions, prompt templates, datasets, etc.) belong to a project.
See [Project Setup](/account-setup/project) for configuration details.
Observability in Freeplay refers to capturing and analyzing the behavior of your AI application through logged data.
### Observability hierarchy
Freeplay uses a three-level hierarchy to organize your AI application logs. From highest to lowest level:
**Session**
The container for a complete user interaction, conversation, or agent run. Sessions group related traces and completions together. Examples include an entire chatbot conversation, a complete agent workflow, or a single user request that triggers multiple LLM calls.
See [Sessions, Traces, and Completions](/core-concepts/observability/sessions-traces-and-completions) for more details.
**Trace**
An optional grouping of related completions and tool calls within a session. Traces represent a functional unit of work, such as a single turn in a conversation, one run of an agent, or a logical step in a multi-step workflow. Traces can be nested to represent sub-agents or complex workflows.
Traces can optionally be given a **name** (like "planning" or "tool\_selection") to represent a specific agent or workflow type. Named traces unlock additional Freeplay features: you can configure evaluation criteria to run against them, create linked datasets for testing, and group similar traces for analysis. See [Agent](#agent) below.
See [Traces](/freeplay-sdk/traces) and [Record Traces](/developer-resources/recipes/record-traces) for implementation details.
**Completion**
The atomic unit of observability in Freeplay. A completion represents a single LLM call, including the input prompt (messages) and the model's response. Every completion is associated with a session and optionally a trace.
See [Recording Completions](/freeplay-sdk/recording-completions) for implementation details.
### Other observability concepts
**Agent**
In Freeplay, an "agent" refers to a named category of traces that represent semantically similar workflows or behaviors. When you give traces the same name (e.g., "research\_agent" or "customer\_support"), Freeplay groups them together as an agent. This grouping enables you to:
* Configure evaluation criteria that run automatically against traces with that name
* Create datasets linked to that agent for testing
* Analyze performance and quality across all traces of that type
It's up to developers to define what constitutes an "agent" in their application. An agent might represent an entire autonomous workflow, a specific sub-task, or any logical grouping that makes sense for your use case.
See [Agents](/practical-guides/agents) for guidance on structuring agent workflows.
**Tool call**
A record of a tool or function call made during an agent workflow, including both the request (tool name and arguments) and the result. Tool calls can be logged in two ways:
* **As part of completions (default):** Tool calls appear in the message history—the tool call in the LLM's output message, and the tool result as an input to the next LLM call. This is simpler and follows standard tool-calling patterns.
* **As explicit tool spans:** Create separate traces with `kind='tool'` for granular visibility into tool execution timing and results as independent spans in the trace view.
See [Tools](/practical-guides/tools) for guidance on recording tool calls.
**Custom metadata**
Contextual information attached to sessions, traces, or completions. Use custom metadata to store data like user IDs, feature flags, business metrics, or workflow identifiers that help with filtering, searching, and analysis.
Custom metadata is recorded via `custom_metadata` fields when creating or updating observability objects. Don't use custom metadata for user feedback like ratings or comments—use [customer feedback](#customer-feedback) instead.
**Customer feedback**
End-user feedback recorded through dedicated feedback endpoints. Customer feedback includes ratings (thumbs up/down, star ratings) and freeform comments. This data receives special treatment in the Freeplay UI due to its distinct utility for quality improvement.
Customer feedback is recorded via the `/completion-feedback/` or `/trace-feedback/` API endpoints.
See [Customer Feedback](/freeplay-sdk/customer-feedback) for implementation details.
**Prompt template**
A versioned configuration that defines everything needed to make an LLM call: the message structure (using Mustache syntax for variables), provider and model selection, request parameters (like temperature), and optionally tool schemas or output structure definitions.
Prompt templates separate the static structure of your prompts from the dynamic variables populated at runtime. This structure enables easy versioning, A/B testing, and dataset creation from production logs.
See [Managing Prompts](/core-concepts/prompt-management/managing-prompts) for more details.
**Environment**
A deployment target for prompt templates, such as `dev`, `staging`, `prod`, or `latest`. Environments let you deploy different versions of your prompts to different stages of your application lifecycle, similar to feature flags.
The `latest` environment always points to the most recently created version of a prompt template. Custom environments can be created for specific use cases.
See [Deployment Environments](/core-concepts/prompt-management/deployment-environments) for configuration details.
**Prompt bundling**
The practice of snapshotting prompt template configurations into your source code repository rather than fetching them from Freeplay's server at runtime. Prompt bundling removes Freeplay from the "hot path" of your application, improving latency and providing compliance benefits for regulated industries.
See [Prompt Bundling](/core-concepts/prompt-management/prompt-bundling) for implementation guidance.
**Mustache**
The templating syntax used in Freeplay prompt templates for variable interpolation. Mustache uses double curly braces (`{{variable_name}}`) for simple substitution and supports conditional logic and iteration.
See [Advanced Prompt Templating Using Mustache](/core-concepts/prompt-management/advanced-prompt-templating-using-mustache) for syntax reference.
**Evaluation**
The process of measuring and scoring the quality of AI outputs. Freeplay supports four types of evaluations:
* **Human evaluation**: Manual review and scoring by team members
* **Model-graded evaluation**: Using an LLM as a judge
* **Code evaluation**: Custom functions that evaluate quantifiable criteria
* **Auto-categorization**: Automated tagging based on specified categories
See [Evaluations](/core-concepts/evaluations/evaluations) for an overview.
**LLM judge**
An LLM-based evaluator that scores AI outputs against specified criteria. Also called "model-graded evaluation." LLM judges can assess nuanced qualities like helpfulness, accuracy, or tone that are difficult to evaluate with code alone.
See [Model-Graded Evaluations](/core-concepts/evaluations/model-graded-evaluations) and [Creating and Aligning Model-Graded Evals](/practical-guides/creating-and-aligning-model-graded-evals) for implementation details.
**Dataset**
A collection of test cases used for evaluation and testing. Datasets can be created by:
* Curating examples from production logs
* Uploading CSV or JSONL files
* Authoring directly in the Freeplay UI
Datasets have schemas that enforce compatibility with specific prompt templates or agents.
The API parameter is `testlist` for legacy reasons, but we use "dataset" in the UI and when referring to this concept in prose.
See [Datasets](/core-concepts/datasets/datasets) for more details.
**Test run**
A batch execution of evaluations against a dataset. Test runs can be:
* **Component-level**: Testing individual prompts or components
* **End-to-end**: Testing complete workflows like full agent runs
Test runs can be initiated via the Freeplay UI or programmatically via SDK for CI/CD integration.
See [Test Runs](/core-concepts/test-runs/test-runs) and [Running a Test Run](/developer-resources/recipes/test-run) for implementation details.
**Review queue**
A workflow for human review of production outputs. Review queues enable teams to:
* Review and annotate completions or traces
* Add structured labels or free-text notes
* Correct LLM judge scores when they're wrong
* Curate examples into datasets
See [Review Queues](/core-concepts/review-queues) for more details.
**Data flywheel**
The continuous improvement cycle enabled by Freeplay's connected workflow. Production logs flow into datasets, which feed evaluations, which inform prompt improvements, which generate better logs. Each iteration strengthens prompts, datasets, evaluation criteria, and testing infrastructure together.
**Freeplay SDK**
Client libraries for integrating Freeplay into your application. Available for:
* **Python**: `freeplay` (install via `pip install freeplay`)
* **TypeScript/Node**: `freeplay` (install via `npm install freeplay`)
* **Java/JVM**: `ai.freeplay:client` (see [SDK Setup](/freeplay-sdk/setup) for Maven/Gradle config)
See [SDK Setup](/freeplay-sdk/setup) for installation and configuration.
**Provider**
The LLM service that processes your prompts. Freeplay supports any provider, including OpenAI, Anthropic, AWS Bedrock, Azure OpenAI, Google, and self-hosted models. The provider is specified in prompt templates or when recording completions.
**Flavor**
The message format used by a specific provider. Different providers expect messages in different formats (e.g., OpenAI's chat format vs. Anthropic's format). Freeplay handles format conversion based on the configured flavor.
## Related resources
* [Why Freeplay?](/getting-started/why-freeplay) - Overview of Freeplay's approach
* [Getting Started](/getting-started/overview) - Quick start guides
* [SDK Documentation](/freeplay-sdk/setup) - Detailed SDK reference
* [API Reference](/openapi/introduction) - HTTP API documentation
# LLMs.txt
Source: https://docs.freeplay.ai/resources/llmstxt
# Support
Source: https://docs.freeplay.ai/resources/support
Get help, check system status, and share feedback.
Our team is here to help. Whether you have a question, need support, or want to share your feedback, we're just an email away.
## Contact us
Reach out to our support team by sending us a message at [support@freeplay.ai](mailto:support@freeplay.ai), and we'll get back to you as soon as possible.
## System status
Check the current status of Freeplay services and subscribe to updates at [status.freeplay.ai](https://status.freeplay.ai).
# Private Deployment (BYOC)
Source: https://docs.freeplay.ai/security-compliance/byoc
Run Freeplay services in your own cloud with Bring-Your-Own-Cloud deployment.
### Reach Out For Access
Interested in BYOC? Access is limited to [Enterprise](https://freeplay.ai/pricing) customers. Please reach out to our team to get access: [privacy@freeplay.ai](mailto:privacy@freeplay.ai)
## BYOC at a glance
Bring-Your-Own-Cloud lets you run **all** Freeplay services inside your own AWS, GCP, or Azure account. You keep full control over data, IAM roles, and network boundaries while Freeplay still updates and supports the software.
**Why choose BYOC?**
* **All prompts and responses remain inside your cloud; Freeplay never receives or stores them.** – Meets strict data-residency or internal-only policies.
* **Reduced vendor-risk reviews** – Auditors see a familiar cloud footprint that you govern.
* **Same product feature velocity** – Freeplay's automated control plane delivers signed updates to the Freeplay application.
* **Minimal DevOps overhead** – when compared to alternative models for privately hosted software, like installing & updating Docker images and setting up & maintaining all of your own serving and database infrastructure
**Architecture Overview**
## BYOC Deep Dive
### How BYOC works
1. **Prepare your environment** – create an empty AWS account, GCP project, or Azure subscription with a resource group.
2. **Deploy the Freeplay agent** – Freeplay generates a bundle for you containing an installer script, Terraform modules, and a Replicated license. The script guides you through choosing a public or private deployment.
3. **Provision infrastructure** – The installer uses [Terraform](https://developer.hashicorp.com/terraform) to provision networking (VPC/VNet), managed PostgreSQL, object storage, KMS encryption, and a Kubernetes cluster. It then deploys Elasticsearch, [NATS](https://nats.io/), and the Freeplay application onto the cluster.
4. **Stay current** – all Freeplay-built Docker images are attested and verified before install. The update mechanism depends on your installation path (see below). We suggest enabling automatic updates via the KOTS admin console to ensure you receive the latest security fixes and features.
```bash theme={null}
# Quick-start (all clouds)
./freeplay_up.sh
```
### Installation paths
Freeplay supports two installation paths for deploying onto your Kubernetes cluster. The right choice depends on your team's operational preferences and is determined during onboarding.
| | Replicated KOTS | Helm CLI (OCI) |
| -------------------- | --------------------------------------------------------------------------------- | --------------------------------------------------------- |
| **Best for** | Teams that prefer a managed admin console with a UI for configuration and updates | Teams that prefer GitOps workflows or direct Helm control |
| **Update mechanism** | Replicated agent polls for releases and applies rolling upgrades automatically | Standard `helm upgrade` against the Freeplay OCI registry |
| **Configuration** | KOTS admin panel or Helm values | Helm values files |
Both paths deploy the same application stack and support the same infrastructure configurations. Freeplay will guide you through the appropriate path during onboarding.
### Prerequisites
Freeplay provides Terraform that provisions everything except the cloud account/project/subscription itself. Some tweaks may apply depending on your networking architecture. Knowing your network configuration in advance helps expedite the process.
#### Tools
All clouds require the following CLI tools on the machine running the installer:
| Tool | Purpose |
| -------------------------------------------------------------- | --------------------------- |
| [Terraform](https://developer.hashicorp.com/terraform/install) | Infrastructure provisioning |
| kubectl | Kubernetes management |
| psql | Database user setup |
| jq | JSON processing |
Plus the cloud-specific CLI for your provider:
| Cloud | Additional tools |
| --------- | ---------------------------------------------------------------- |
| **AWS** | `aws` CLI, `session-manager-plugin` (for bastion access via SSM) |
| **GCP** | `gcloud` CLI, `cloud-sql-proxy` |
| **Azure** | `az` CLI, `kubelogin` |
#### Infrastructure requirements
| Item | Requirement |
| --------------- | ---------------------------------------------------------------------------------------------------------------- |
| Cloud project | Empty AWS account / GCP project / Azure subscription + resource group |
| Kubernetes | 1.32+ (EKS, GKE, AKS) |
| Database | Managed PostgreSQL 17 — RDS, Cloud SQL, or Azure PostgreSQL Flexible Server |
| Object storage | S3 / GCS / Azure Blob Storage (two buckets: assets and data export) |
| IAM | Deployment user/role with permissions described below; least-privilege workload roles are generated by Terraform |
| Networking | VPC/VNet with private subnets and NAT for egress; optional VPC/VNet peering or Private Service Connect |
| Outbound egress | Replicated, WorkOS, and optionally Datadog and Mixpanel (HTTPS only, see domains below) |
### What gets deployed
The following services run inside your Kubernetes cluster. Default values are shown — all replica counts, resource limits, and storage sizes are configurable via the KOTS admin panel or Helm values at initial deploy and on subsequent updates.
| Component | Type | Details |
| --------------------- | ---------------- | ------------------------------------------------------------------------ |
| **Freeplay web app** | Deployment (HPA) | Main application; scales 2–10 replicas by default |
| **Index Completions** | Deployment | Search indexer consuming from NATS, writing to Elasticsearch; 2 replicas |
| **Data Export** | CronJob | Exports event data from NATS to object storage |
| **Elasticsearch** | ECK-managed | Search backend; 3 replicas, 200 Gi storage per pod |
| **NATS** | JetStream | Message broker; 5 replicas, 10 Gi storage per pod |
Infrastructure provisioned outside the cluster:
| Component | AWS | GCP | Azure |
| ------------------ | ----------------- | ----------------------- | ----------------------------- |
| **Database** | RDS PostgreSQL 17 | Cloud SQL PostgreSQL 17 | PostgreSQL Flexible Server 17 |
| **Object storage** | S3 | GCS | Blob Storage |
| **Secrets** | Secrets Manager | Secret Manager | Key Vault |
| **Encryption** | KMS | Cloud KMS | Key Vault |
| **DNS** | Route 53 | Cloud DNS | Azure DNS |
| **TLS** | ACM | Managed SSL Certificate | cert-manager / App Gateway |
#### AWS Services and IAM
AWS services are available by default in your account. The deployment primarily requires appropriate IAM permissions.
```
# AWS Services Used
## Compute & Containers
EKS (Elastic Kubernetes Service) — Auto Mode
EC2 (for NAT Gateway, SSM bastion)
## Database
RDS (PostgreSQL 17, default db.r6i.2xlarge, Multi-AZ)
## Storage
S3 (assets bucket and data-export bucket, KMS-encrypted)
## Security & Secrets
Secrets Manager (DB credentials, connection strings)
KMS (encryption at rest for S3, RDS, secrets)
IAM (IRSA for pod-level permissions)
## Networking
VPC, private and public subnets (2 AZs), NAT Gateway
Security Groups (RDS, bastion, EKS)
Optional VPC peering (cross-account, cross-region)
## DNS & Certificates
Route 53 (hosted zone)
ACM (Certificate Manager, DNS-validated)
## Monitoring
CloudWatch Logs
## Optional
Lambda (isolated-VPC code evaluation)
Bedrock IAM role (for model access)
```
```
# IAM — Deployment user/role
# Broad permissions needed during initial provisioning:
eks:*
ec2:*
rds:*
s3:*
secretsmanager:*
kms:*
iam:* (for creating IRSA roles and policies)
route53:*
acm:*
elasticloadbalancing:*
logs:*
lambda:* (if code evals enabled)
sts:AssumeRole (if Bedrock or cross-account peering)
# IAM — Pod workload identity (IRSA)
s3:PutObject, s3:GetObject, s3:ListBucket (on Freeplay buckets)
secretsmanager:GetSecretValue, secretsmanager:DescribeSecret (on DB secrets)
kms:DescribeKey, kms:Encrypt, kms:Decrypt, kms:ReEncrypt*,
kms:GenerateDataKey, kms:GenerateDataKeyWithoutPlaintext
# IAM — External-DNS (IRSA)
route53:ChangeResourceRecordSets
route53:ListHostedZones, route53:ListResourceRecordSets
# IAM — Bedrock (optional)
bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream
sts:AssumeRole (to assume Bedrock role)
```
#### Google Cloud APIs and IAM
The `./freeplay_up.sh` script validates required APIs and the Terraform modules enable them automatically during provisioning. We recommend keeping resources isolated by project.
```
# Required Google Cloud APIs
## Compute & Containers
compute.googleapis.com
container.googleapis.com
## IAM & Workload Identity
iam.googleapis.com
## Database
sqladmin.googleapis.com
## Secrets & Encryption
secretmanager.googleapis.com
cloudkms.googleapis.com
## DNS & TLS
dns.googleapis.com
certificatemanager.googleapis.com
servicenetworking.googleapis.com
## Storage
storage.googleapis.com
## AI / LLM Access
aiplatform.googleapis.com
generativelanguage.googleapis.com
## Networking (peering / PSC scenarios)
networksecurity.googleapis.com
networkmanagement.googleapis.com
servicedirectory.googleapis.com
## Code Evaluation (optional)
cloudbuild.googleapis.com
artifactregistry.googleapis.com
cloudfunctions.googleapis.com
run.googleapis.com
vpcaccess.googleapis.com
```
```
# IAM — Terraform service account (deployment)
roles/editor
roles/compute.networkAdmin
roles/cloudkms.admin
roles/secretmanager.admin
roles/cloudsql.admin
roles/container.admin
roles/storage.admin
roles/iam.serviceAccountAdmin
roles/resourcemanager.projectIamAdmin
roles/serviceusage.serviceUsageAdmin
# The user running the installer also needs:
roles/iam.serviceAccountTokenCreator (on the Terraform service account)
# IAM — GKE node service accountroles/logging.logWriter
roles/monitoring.metricWriter
roles/monitoring.viewer
roles/stackdriver.resourceMetadata.writer
roles/storage.objectViewer
# IAM — Freeplay app workload identityroles/monitoring.metricWriter
roles/storage.objectAdmin
roles/logging.logWriter
roles/cloudsql.client
# IAM — External-DNS workload identityroles/dns.admin
```
#### Azure Resource Providers and Roles
These resource providers must be registered in your Azure subscription. The `./freeplay_up.sh` script checks for and registers them before running Terraform.
```
# Resource Providers that must be registered
## Compute & Containers
Microsoft.ContainerService
Microsoft.Compute
## Networking
Microsoft.Network
## Secrets & Encryption
Microsoft.KeyVault
## Storage
Microsoft.Storage
## Database
Microsoft.DBforPostgreSQL
## Identity
Microsoft.ManagedIdentity
## Monitoring
Microsoft.OperationalInsights
```
```
# Azure Roles
## For the deployment service principal or user
Contributor (on resource group)
User Access Administrator (on resource group)
## For the AKS workload identity (user-assigned managed identity)
Key Vault access policy: Get/List secrets, Get/List/Encrypt/Decrypt/WrapKey/UnwrapKey keys
Storage Blob Data Contributor (on storage accounts)
DNS Zone Contributor (on DNS zone)
Network Contributor (on VNet)
## For AKS cluster admins
Azure Kubernetes Service RBAC Cluster Admin (on AKS cluster)
```
### Security & compliance
* **Data residency** – all sensitive customer data including prompts, responses, and evaluations stay in your cloud account. *Optional* ability to share support bundles for troubleshooting scenarios.
* **Secrets** – stored in your cloud KMS-backed secret manager (AWS Secrets Manager, GCP Secret Manager, or Azure Key Vault); never transmitted to Freeplay.
* **Encryption** – data at rest is encrypted with cloud KMS keys provisioned in your account. Keys rotate automatically every 90 days on all three clouds.
* **Network isolation** – databases are deployed on private subnets only, with no public IP. Kubernetes API servers can be made private with VPC/VNet peering.
* **Least-privilege IAM** – Terraform creates workload-identity bindings (IRSA, GCP Workload Identity, Azure Managed Identity) scoped to only the resources each service needs.
* **Updates** – all Freeplay-built Docker images are attested and verified before install; only metadata (version, health ping) is sent to the control plane.
* **Configurable defaults** – the Terraform modules expose variables for database backups and retention, high-availability mode, deletion protection, VPC flow logs, network policies, and more. Defaults are production-ready, but every knob is tunable to match your security and compliance requirements.
### Costs
* **Freeplay subscription** – BYOC is available only for [Enterprise-tier contracts](https://freeplay.ai/pricing). Please contact Sales for more info about access.
* **Your cloud** – you pay for all compute nodes, database, Elasticsearch storage, object storage, and egress. Typical mid-sized install runs approximately $1.2k–$2k/mo in AWS us-east-1.
### BYOC Outbound Egress Requirements
All outbound traffic uses **port 443 / HTTPS (TLS 1.2+)**.
#### 1. WorkOS Authentication (Required)
| | |
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Domains** | `*.authkit.app`, `*.workos.com` |
| **Purpose** | User authentication and authorization |
| **Data transmitted** | User email |
| **IP ranges** | Cloudflare published ranges — see [WorkOS On-Prem Deployment](https://workos.com/docs/on-prem-deployment/introduction) and [Cloudflare IPs](https://www.cloudflare.com/ips/) |
#### 2. Replicated (Conditional)
| | |
| -------------------- | ----------------------------------------------------------------------------------------- |
| **Domains** | `*.replicated.com`, `*.replicated.app` |
| **Purpose** | Agent polls for signed application & infrastructure releases |
| **Data transmitted** | Agent ID, current version, signed artifacts |
| **IP ranges** | [replicatedhq/ips](https://github.com/replicatedhq/ips/blob/main/ip_addresses.json) |
| **Firewall rules** | [Replicated Customer Firewalls](https://community.replicated.com/t/customer-firewalls/55) |
#### 3. Datadog (Conditional)
| | |
| -------------------- | -------------------------------------------------------------------------------------------------------- |
| **Domains** | `*.datadoghq.com` (regional endpoints) |
| **Purpose** | Metrics and traces for 24×7 ops SLA |
| **Data transmitted** | Health metrics, cluster stats, anonymous performance metrics, scrubbed logs (no prompt/response content) |
| **IP ranges** | [ip-ranges.datadoghq.com](https://ip-ranges.datadoghq.com/) |
#### 4. LLM Provider Endpoints (Conditional)
| | |
| -------------------- | ---------------------------------------------------------------------------------------------------------- |
| **Domains** | Provider-specific (e.g., `api.openai.com`, `bedrock-runtime.*`, `*.anthropic.com`, Azure OpenAI endpoints) |
| **Purpose** | Inference calls initiated by Freeplay application |
| **Data transmitted** | Prompts, context, parameters, model responses |
| **Required?** | Yes, if you use external hosted models |
| **How to restrict** | Use VPC endpoints or Private Link where the provider offers them |
#### 5. Mixpanel (Optional)
| | |
| -------------------- | ----------------------------------------------------------------------- |
| **Domains** | `*.mixpanel.com` |
| **Purpose** | Product-usage analytics (UX insights) |
| **Data transmitted** | Anonymous instance ID, UI events (no sensitive prompt/response data) |
| **How to disable** | Set the Mixpanel token to empty via the KOTS admin panel or Helm values |
Destinations 1 and 2 are always required for a supported install. 3 is strongly recommended. Destination 4 is required only when your workloads invoke external hosted models — BYOC also supports running local/on-site models if you need zero data egress for LLM traffic.
**Key Points**
* You decide which LLM providers and regions to enable in the Freeplay UI.
* All calls use HTTPS (TLS 1.2+). AWS enforces TLS 1.3 via `ELBSecurityPolicy-TLS13-1-2-PQ-2025-09`. GCP enforces a RESTRICTED SSL policy. Azure supports TLS 1.2 and TLS 1.3.
* Freeplay's log policy prevents user content from landing in Datadog.
* If you need a *completely* offline install, talk to us. An artifact-mirror and on-prem observability stack are on the roadmap.
## FAQ
| Question | Answer |
| ----------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Can we air-gap entirely? | Not yet. The Replicated agent needs periodic egress for updates and license checks. |
| Is bring-your-own KMS supported? | Yes. All encryption keys are stored on a KMS in your account (AWS KMS, GCP Cloud KMS, or Azure Key Vault). GCP supports configurable protection levels (SOFTWARE or HSM). |
| What about bring-your-own VPC or cluster? | Yes. Terraform variables let you supply an existing VPC/VNet (`create_vpc = false`) and an existing Kubernetes cluster (`create_eks_cluster`, `create_gke_cluster`, or `create_aks_cluster = false`). You can also bring your own DNS zone and security groups. While a vanilla deployment is easiest, BYOC adapts to pre-established networking rules. |
| Can I scale horizontally? | The Helm values file exposes replicas and resource limits. The Freeplay web app uses a Horizontal Pod Autoscaler (default 2–10 replicas, scaling on CPU and memory). Elasticsearch (3 replicas) and NATS (5 replicas) are also configurable. |
| Private-only access? | Yes. The installer script asks whether you want public or internal-only load balancing. For private access, you can use VPC/VNet peering or Private Service Connect (GCP). When peering is enabled, GCP additionally provisions Cloud Armor rules to restrict traffic to known CIDR ranges. |
| Is the site reachable from the public Internet? | We can provision either a public internet site, or create a private site that is only resolvable or firewalled to your virtual network or connection. |
| What persistent storage is part of this deployment? | Managed PostgreSQL (RDS / Cloud SQL / Azure PostgreSQL Flexible Server) for the source of truth. S3 / GCS / Blob Storage for multi-modal artifacts and data export. Block storage (EBS / Persistent Disk / Azure Managed Disk) for Elasticsearch indices (default 200 Gi per pod) and NATS JetStream (default 10 Gi per pod). Storage sizes are configurable via the KOTS admin panel or Helm values. |
| I've heard Kubernetes can be unreliable with persistent volumes. Does that put my data at risk? | No. BYOC keeps the source-of-truth in a managed database service (RDS, Cloud SQL, or Azure PostgreSQL Flexible Server) that runs outside Kubernetes. Everything on in-cluster volumes (search indices, NATS JetStream, temporary files) can be rebuilt from that database. If a node or the whole cluster is lost, you spin up new nodes, redeploy the chart, and point it at the same database. The application comes back with no data loss. |
| Can I use an existing Elasticsearch instance? | Yes. Set `elasticSearch.enabled: false` in the Helm values and provide your external host via `freeplay.ELASTIC_HTTP_HOST`, bypassing the in-cluster ECK deployment. |
| Is there a Terraform backend? | Yes. State is stored in an S3 bucket (or GCS / Azure Blob equivalent) in your cloud account, so the Terraform state stays within your security perimeter. |
| What about code evaluation? | Code evaluation is an optional feature available on AWS (Lambda in an isolated VPC with no egress) and GCP (Cloud Functions with Private Service Connect, egress-deny firewall rules, and DNS blocking). Contact us for details. |
| Can I preview the Terraform and networking rules / IAM policies? | Yes, please contact us. |
# Data Retention
Source: https://docs.freeplay.ai/security-compliance/data-retention-policy
Freeplay handles different types of data in specific ways to balance your needs for analysis with privacy and storage considerations. This page explains how long we keep your data and why.
## How Data Retention Works
When you use Freeplay, we collect various types of data through the SDK and within the application. Here's what you need to know about how long this data is kept.
### Logged Data (Standard Retention)
Most sessions, traces, and completions recorded through the Freeplay SDK are kept for **90 days** from the date of recording by default. This gives you enough time to analyze results and troubleshoot issues without storing potentially sensitive interaction data indefinitely.
After 90 days, this data is automatically deleted from our systems by default. For customers on our Enterprise plans, the standard retention window is configurable.
### Logged Data (Extended Retention)
Some data is automatically kept for longer periods because it's referenced elsewhere in the Freeplay platform. You can always remove it yourself, but we'll keep your data **indefinitely** if it meets any of these criteria:
* It's part of a Test Run
* It's included in a Dataset
* It has Customer Feedback logged against it from your application
* It's assigned to a Review Queue
* It has associated review information (e.g., "in progress," "completed")
* It has evaluations or labels applied to it
This extended retention helps you maintain test cases, evaluate models over time, and track the quality of your systems. This data is kept until you explicitly delete it through the Freeplay platform.
If needed, the deletion can be done via the API by fetching all of the sessions using the [list sessions](https://docs.freeplay.ai/api-reference/search-&-analytics/list-sessions) endpoint and then deleting the session by the [delete session endpoint](https://docs.freeplay.ai/api-reference/observability/delete-session).
### Freeplay Application Data
All your configurations and settings within Freeplay are kept indefinitely, including:
* Prompt Templates
* Model Settings
* Review Queues
* Test Runs
* Comparisons
* Authorization Settings (user accounts, roles, API keys)
* Project Configurations
This ensures your Freeplay environment remains stable and functional over time. You can delete this data at any time through the platform.
### Custom Retention Options
If you need to keep standard data beyond the 90-day period, we offer custom retention options.
Contact our team to discuss custom retention options for your specific use case.
## What Happens When You Close Your Account
If you decide to close your Freeplay account, here's what you can expect:
* We'll notify you before account termination
* You should export any data you want to keep (we provide tools to help with this)
* After a 30-day grace period, we will delete all your remaining data
* We'll only keep information required by law or for legitimate business purposes (like billing records)
## Questions About Data Retention?
If you have any questions about how we handle your data or need help with custom retention options, please contact Freeplay [support@freeplay.ai](mailto:support@freeplay.ai). We're here to help!
***
*Note: We may update our data retention practices as technology or legal requirements change. We'll notify you of any significant changes to how we handle your data.*
# GDPR
Source: https://docs.freeplay.ai/security-compliance/general-data-protection-regulation-gdpr
At Freeplay, the privacy and security of your customer data is one of our top priorities. General Data Protection Regulation (GDPR) applies not only to EU-based businesses, but also to any business that controls or processes data of EU citizens. Our entire organization is hard at work ensuring that our practices are GDPR-compliant.
| **Section** | **Explanation** |
| ------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Individual in Charge of GDPR** | Eric Ryan, CTO ([privacy@freeplay.ai](mailto:privacy@freeplay.ai)) |
| **Data Protection Officer** | Eric Ryan, CTO ([privacy@freeplay.ai](mailto:privacy@freeplay.ai)) |
| **Purpose of Processing** | Freeplay provides a platform for testing, evaluating, monitoring, and optimizing AI-driven products. We process LLM request and response data to help teams improve their generative AI features. These requests may at times include end user data from our customers’ AI systems. Freeplay processes this data for purposes including assisting our customers with product experimentation, automated testing, LLM observability, and data labeling & curation. The use of the Freeplay application also includes data related to our customers' authorized users of Freeplay, such as names, email addresses, and user IDs. We may engage third-party service providers to assist in providing our services, always in accordance with our data processing agreements and with a focus on data privacy and security. |
| **Lawful Basis of Processing and Consent** | Under Article 6 of GDPR, our processing falls under:- **Consent**: Via [Privacy Policy](https://freeplay.ai/privacy) . Removal of consent can be requested by contacting [privacy@freeplay.ai](mailto:privacy@freeplay.ai). |
* **Contract**: Via [Terms of Service](https://freeplay.ai/legal/tos) or Master Services Agreement with business customers, which gives Freeplay permission to manage the Personal Data of our customers' employees. |
\| **Withdrawal of Consent (or Opt-Out)** | For end users, withdrawal of consent or opting out after initial consent/opt-in is available by emailing [privacy@freeplay.ai](mailto:privacy@freeplay.ai)\ |
\| **Deletion Policy** | Deletion of Freeplay Customer Data upon termination, cancellation, or expiration of the agreement. Our customers can also delete any of their end user data via the Freeplay web application or API, in order to honor their own deletion obligations. Data deletion for website visitors can be requested by contacting [privacy@freeplay.ai](mailto:privacy@freeplay.ai) . |
\| **Data Access / Modification / Portability** | Users can access, modify, and delete their data directly from the Freeplay web application. Website visitors can request a copy or update their data by emailing [privacy@freeplay.ai](mailto:privacy@freeplay.ai) . |
\| **Data Protection Information** | Freeplay deploys and maintains industry-standard best practices attested to in a [SOC 2 Type 2 report](https://trust.freeplay.ai/) covering security, confidentiality, and availability. Further information is contained in Freeplay's Terms of Service and Data Processing Addendum. |
\| **Notification of Breach** | Freeplay’s breach notification process is outlined within our Terms of Service, Data Processing Addendum, and Incident Response Policy, which is available upon request. |
\| **EU Representative Contact** | Osano International Compliance Services Limited ATTN: SLCA 3 Dublin Landings North Wall Quay Dublin 1 D01 C4E0 Ireland |
\| **UK Representative Contact** | Osano UK Compliance Ltd ATTN: SLCA 42 – 46 Fountain Street Belfast Antrim BT1 5EF United Kingdom |
# Compliance with Prompt Bundling
Source: https://docs.freeplay.ai/security-compliance/production-prompt-bundling-compliance-guard-rails
Use Prompt Bundling plus a lightweight GitHub Actions check to ensure every production prompt change is peer‑reviewed and fully auditable—meeting common controls such as SOC 2 CC4.1, ISO 27001 A.14.2.5, and PCI‑DSS 6.4.5.
## Why it matters
* Prompt text *is* business logic. Untracked changes can introduce risk or inconsistent behavior.
* Most compliance frameworks require that **production code** (including prompts) is:
* **Immutable** after deployment
* **Peer‑reviewed** before release
* **Traceable** with a full audit trail
* Bundling a `prod` **Prompt Template** into an artifact that lives in Git provides these guarantees with almost zero additional tooling.
* This also protects your application in the unlikely event that you're application cannot connect to the Freeplay platform
***
## Prerequisites for this guide
* GitHub repository
* GitHub Actions enabled
* Freeplay CLI
* A Prompt Template already promoted to the **`prod`environment**
***
## Step 1 · Bundle the production prompt
Run during your release pipeline:
```bash theme={null}
freeplay download \
--environment prod \
--output-dir bundled_prompts \
--project-id
```
This writes an **immutable** JSON artifact containing:
* Prompt text
* Model/provider selection
* All request parameters
***
## Step 2 · Open a pull request if the bundle changes
Add the following workflow file to `.github/workflows/prompt-bundle-guard.yml`:
```yaml yaml theme={null}
name: Update Bundled Prompts
on:
push:
branches: [main]
jobs:
prompt-bundle-guard:
runs-on: ubuntu-latest
permissions:
contents: write # push temp branch
pull-requests: write # open PR
steps:
- uses: actions/checkout@v4
- name: Install Freeplay CLI
run: python -m pip install --upgrade freeplay
- name: Re‑bundle prod prompts
run: |
freeplay download \
--environment prod \
--output-dir bundled_prompts \
--project-id
- name: Open PR if bundle changed
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
if git diff --quiet --exit-code -- bundled_prompts; then
echo "No prompt updates detected." && exit 0
fi
git config user.name "Freeplay Prompt Bundler"
git config user.email "[privacy@freeplay.ai]"
branch="auto/prompt-bundle-$(date +%s)"
git switch -c "$branch"
git add bundled_prompts
git commit -m "chore(prompts): update prod bundled prompt"
git push --set-upstream origin "$branch"
gh pr create \
--title "Update production bundled prompts" \
--body "Automated PR generated by Prompt Bundler workflow." \
--label "prompt-bundle" \
--base main \
--head "$branch" \
--draft
```
### What happens?
1. **Bundle** – Regenerates the prod bundle.
2. **Diff** – If anything changed inside `bundled_prompts/`, the workflow:
* Pushes the change on a new `auto/prompt-bundle-*` branch.
* Creates a **draft PR** that a human must review & merge.
3. Once merged, the new bundle is locked in Git history.
> **Tip:** Protect `main` with “Require PR approval” to enforce peer review.
***
## Step 3 · Pin the Bundled Prompt for production systems at runtime
```python python theme={null}
from pathlib import Path
from freeplay import Freeplay
from freeplay.resources.prompts import FilesystemTemplateResolver
# Point the resolver at your committed bundle directory
fp_client = Freeplay(
freeplay_api_key=freeplay_key,
api_base=freeplay_api_base,
template_resolver=FilesystemTemplateResolver(Path("bundled_prompts"))
)
# Invoke the production template as usual
```
The SDK reads the **local** bundle, so the application can only execute prompts that passed the PR gate.
***
## Compliance mapping
| Framework | Requirement | How the workflow satisfies it |
| ------------------------------- | --------------------------------------------------- | ----------------------------------------------------- |
| **SOC 2 CC4.1** | Peer review of production changes | PR approval required before merge |
| **ISO 27001 A.14.2.5** | Secure engineering principles & immutable artifacts | Bundled Prompt is hashed & version‑controlled |
| **PCI‑DSS 6.4.5** | Formal approval prior to production | PR review & protected branch policies |
| **HIPAA §164.308(a)(1)(ii)(D)** | System activity review & audit trails | Git + GitHub Actions logs show who changed what, when |
***
## FAQ & Troubleshooting
**Q : What if we maintain multiple prod environments (e.g., per‑tenant)?** A : Run `freeplay download --env prod-` (or similar, given your environment naming conventions) for each environment and store each bundle under its own path.
**Q : Can I use Bitbucket Pipelines or GitLab CI instead?** A : Yes mirror the same logic: re‑bundle, diff, and open a merge request when changes are detected.
**Q : How do I invalidate a bad prompt quickly?** A : Revert the bundle commit or promote a previous prompt template version in the Freeplay dashboard and re‑run the workflow.
# Overview
Source: https://docs.freeplay.ai/security-compliance/security-overview
At Freeplay, security is a top priority. We've implemented robust measures to protect your data and ensure the safety of our platform.
## Key Management
Freeplay uses advanced cryptographic key management processes to secure sensitive customer data and generate API keys for its private API. Our key management approach includes:
### Cryptographic Encryption Keys
1. **Key Generation and Distribution:** We use [Google Cloud Key Management](https://cloud.google.com/security/products/security-key-management?hl=en) for generating symmetric encryption keys. These keys are created using the "Google symmetric key" encryption algorithm. Each key is unique to the Freeplay customer and is local to the region where the customer's Google Cloud Run application is deployed.
2. **Key Storage:** Our encryption keys are stored and managed according to the [FIPS 140-2 Level 1 standard](https://en.wikipedia.org/wiki/FIPS_140-2). We use Google Cloud Key Management to store keys securely, ensuring they are encrypted both at rest and in transit. Access to these keys is restricted to the Cloud Run service account running the Freeplay application, adhering to the principle of least privilege.
3. **Key Rotation:** To enhance security and mitigate the risk of key compromise, we rotate encryption keys every 90 days. This process is automated and managed through Google Cloud Key Management, ensuring seamless updates without service interruption.
4. **Accountability and Audit:** All encryption and decryption actions are logged using Google Cloud Audit logging. These logs provide an audit trail and are only accessed in case of an incident or investigation, ensuring traceability and accountability for all key management actions.
### Freeplay API Keys
For access to the Freeplay API, we generate and validate API keys using the following process:
1. **Generation:** API keys are created using a cryptographically secure randomly generated string.
2. **Display and Storage:** The full API key is displayed to the requesting customer only immediately after generation. After this initial display, we use the Argon2 algorithm to create a one-way hash of the key before storage. For user convenience, we persist only the last 4 characters of the original key.
3. **Validation:** When incoming API requests occur, customer API keys are hashed and validated against the stored hash to allow or deny access.
4. **Management:** Customers can revoke or rotate their API keys at any time using the Freeplay web application.
## Access Control
At Freeplay, we implement strict access control measures to ensure the security and integrity of our systems and your data:
1. **Principle of Least Privilege:** We adhere to the principle of least privilege, which means that users and systems are granted the minimum levels of access – or permissions – needed to perform their functions. This minimizes the potential impact of any security breach.
2. **Regular Access Reviews:** We conduct regular reviews of access rights to ensure that permissions remain appropriate as roles change within our organization.
3. **Audit Logging:** All sensitive actions, such as changes to system configurations or modifications to access rights are logged and monitored.
## Vulnerability Management
At Freeplay, we take a proactive approach to vulnerability management:
1. **Continuous Scanning:** We employ advanced tools to conduct continuous vulnerability scanning across our entire infrastructure. This allows us to identify potential weaknesses in real-time.
2. **Regular Penetration Testing:** We engage third-party security experts to perform regular penetration testing. These tests simulate real-world attack scenarios, helping us uncover and address vulnerabilities that automated scans might miss.
3. **SLA-Driven Resolution:** Our team adhere to a Service Level Agreement (SLA) for addressing vulnerabilities.
## Data Isolation
We've implemented tenant isolation at the database layer using row-level security. This ensures that each customer's data remains strictly segregated, preventing any unauthorized access or data leakage between different tenants sharing our infrastructure.
## Multi-Factor Authentication Requirement
We require Multi-Factor Authentication (MFA) for all users of Freeplay systems. This is not optional, as it adds an extra layer of protection to your account.
## Certifications & Compliance
Freeplay maintains industry-standard certifications to meet enterprise security requirements. View our certification details at [trust.freeplay.ai](https://trust.freeplay.ai).
* **SOC 2 Type II:** Freeplay has achieved SOC 2 Type II certification, demonstrating our commitment to security, availability, processing integrity, confidentiality, and privacy.
* **HIPAA:** Freeplay supports HIPAA compliance for healthcare organizations. We offer Business Associate Agreements (BAAs) for both our multi-tenant cloud platform and [BYOC (Bring Your Own Cloud)](/security-compliance/byoc) deployments. Contact our sales team to execute a BAA for your organization.
* **GDPR:** Freeplay is fully compliant with the General Data Protection Regulation. [Learn more about our GDPR compliance](/security-compliance/general-data-protection-regulation-gdpr).
## Private Hosting
For customers with heightened security requirements, we offer a private hosting solution. This uses a single tenant deployment along with a secure and highly available site-to-site VPN connection. It allows you to keep your data within your own infrastructure, providing an additional layer of control and security. [Learn more about our private hosting option](/security-compliance/byoc).