Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions apps/docs/components/icons.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -9936,3 +9936,16 @@ export function PitchBookIcon(props: SVGProps<SVGSVGElement>) {
</svg>
)
}

/** TypeSafe’s official mark from https://typesafe.ai. */
export function TypeSafeIcon(props: SVGProps<SVGSVGElement>) {
return (
<svg {...props} viewBox='0 0 21 32' fill='currentColor' xmlns='http://www.w3.org/2000/svg'>
<path
fillRule='evenodd'
clipRule='evenodd'
d='M10.36 0.071C10.55 -0.054 10.707 -0.008 10.757 0.166L15.57 2.792L15.33 3.23L15.748 2.953V8.81L20.542 11.425L20.224 12.008L20.742 11.663V23.613C20.742 23.889 20.556 24.238 20.326 24.391L10.383 31.02V30.402L10.12 30.885L5.323 28.269C5.124 28.365 4.972 28.262 4.972 28.012V22.153L0.398 19.659C0.176 19.797 0 19.697 0 19.428V7.478C0 7.201 0.186 6.853 0.416 6.7L10.36 0.071ZM6.27 27.645L10.49 29.947L19.434 23.985L15.212 21.684L6.27 27.645ZM10.775 12.743V18.218C10.775 18.494 10.589 18.842 10.359 18.995L5.804 22.032V26.957L14.915 20.882V9.983L10.775 12.743ZM15.747 20.827C15.747 20.83 15.746 20.832 15.746 20.835L19.91 23.107V12.219L15.747 9.948V20.827ZM1.34 19.033L5.401 21.249L9.435 18.56L5.373 16.345L1.34 19.033ZM0.832 7.423V18.373L4.972 15.613V10.139C4.972 9.863 5.158 9.515 5.388 9.361L9.944 6.324V1.349L0.832 7.423ZM5.804 15.44L9.943 17.697V12.913L5.804 10.655V15.44ZM6.271 9.771L10.335 11.988C10.343 11.982 10.351 11.975 10.359 11.97L14.374 9.291L10.313 7.077L6.271 9.771ZM10.776 6.189L14.916 8.447V3.573L10.776 1.316V6.189Z'
/>
</svg>
)
}
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ import { Callout } from 'fumadocs-ui/components/callout'
| `MISTRAL_API_KEY` | Mistral |
| `XAI_API_KEY_1` | xAI |
| `KIMI_API_KEY_1` | Moonshot Kimi |
| `TYPESAFE_API_KEY_1` / `_2` / `_3` | TypeSafe Jev hosted key rotation |
| `ZAI_API_KEY_1` | Z.ai |
| `TOGETHER_API_KEY` | Together AI |
| `FIREWORKS_API_KEY` | Fireworks AI |
Expand Down
26 changes: 25 additions & 1 deletion apps/docs/content/docs/workflows/blocks/agent.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,30 @@ For a custom cloud deployment, enter its provider prefix and model ID: `azure/my

Ollama Cloud, OpenRouter, Fireworks, Together AI, Baseten, Ollama, vLLM, and LiteLLM load their available models from the configured provider. New models appear through that discovery without a Sim catalog release. You can also enter a namespaced ID directly, such as `ollama-cloud/deepseek-v4.1-flash`, `openrouter/provider/model`, or `ollama/my-local-model`. Provider prefixes are case-insensitive; the model ID after the prefix keeps its original casing.

### Jev evaluation models

Select `jev-latest` from TypeSafe in the Agent model selector. Hosted Sim supplies a key and bills model usage through the normal credit system; workspace or organization BYOK keys override the hosted key without model charges. Self-hosted users enter their TypeSafe key in the block. Use `jev-1.13.0` to pin a version or `jev-preview` to follow preview releases. These models use **State** and **Questions** in place of conversational messages. State accepts text or a reference to a JSON object or array. Questions is a JSON object keyed by the answer names you want:

```json
{
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": { "billing": "Payments and invoices", "support": "Product issues" }
},
"urgency": {
"type": "score",
"instructions": "How urgent is the request?",
"criteria": ["Routine", "Soon", "Immediate"]
},
"resolved": { "type": "noul", "instructions": "Has the request been resolved?" }
}
```

Read results from `<agent.answers>` or expand an answer in the reference picker, such as `<agent.answers.route.choice>`. Each Choice answer includes `choice`, `probabilities`, and `confidence`; each Score answer includes `score`, `legend`, `probabilities`, and `confidence`; each Noul answer includes `noul`, a probability from 0 to 1. `content` contains the same answers as JSON text, and the standard model, token, timing, and cost outputs remain available. Use a Condition block to route on these results.

Jev evaluates the supplied state in one request. Chat messages, files, tools, skills, conversation memory, response-format schemas, and chat model fallbacks are hidden for these models. Saved settings return when you switch back to a chat model. TypeSafe documents a 64,000-token total request limit and a 32,000-token limit for state plus the longest question. See [TypeSafe's model documentation](https://docs.typesafe.ai/models) and [question formats](https://docs.typesafe.ai/api).

### Files

Files for the model to read: images for a vision-capable model, or documents for text. Upload them on the block, or pass a file from an earlier block, such as an upload trigger or an [API](/workflows/blocks/api) response, with a connection tag.
Expand Down Expand Up @@ -104,7 +128,7 @@ Some settings live under advanced, or appear only for models that support them:
- **Max output tokens.** Caps the response length. Defaults to the model's full limit.
- **Reasoning effort / Thinking level.** For models with extended reasoning, how much the model thinks before answering. Higher is more thorough but slower and costs more tokens.
- **Prompt caching.** For Anthropic Claude models, reuses the system prompt and tool definitions between runs instead of re-reading them every time. Cached input costs a tenth of the normal rate, but writing the cache costs 1.25x, so leave it off for one-off runs and turn it on when the same agent runs repeatedly. The cache covers a prefix only if it reaches 1,024 tokens (2,048 on Haiku) — below that Anthropic ignores it and nothing changes. Entries expire after five minutes of no use.
- **API key.** Your key for the chosen provider. Hidden on hosted Sim, which supplies one.
- **API key.** Your key for the chosen provider. Hidden when hosted Sim supplies a key for the selected model, including Jev.
- **Fallback models.** An ordered list of up to five models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, once each, and `<agent.model>` reports the model that answered. On hosted Sim, hosted models use your workspace's BYOK or platform credentials; local and self-hosted installations may still require a key. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. A stored row key stops applying when its key field is hidden. Providers that require family-specific credentials, such as Vertex, can only be fallbacks for a selected model of the same family. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it; leave it empty and the provider's default applies.
- **Retry on fail.** Retries the selected model after a failure, up to a maximum number of tries with a wait between them. When its tries run out, the fallback models are tried in order, once each, with no wait before the first of them. A fallback is never retried. See [Retries and fallbacks](#retries-and-fallbacks) for how recorded tool results are reused and when a tool can execute again.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@ import {
SerperIcon,
TinyFishIcon,
TogetherIcon,
TypeSafeIcon,
WizaIcon,
xAIIcon,
ZaiIcon,
Expand Down Expand Up @@ -130,6 +131,13 @@ const PROVIDERS: (BYOKManagerProvider & { id: BYOKProviderId })[] = [
description: 'LLM calls',
placeholder: 'sk-...',
},
{
id: 'typesafe',
name: 'TypeSafe',
icon: TypeSafeIcon,
description: 'Jev evaluation models',
placeholder: 'Enter your TypeSafe API key',
},
{
id: 'fireworks',
name: 'Fireworks',
Expand Down Expand Up @@ -352,6 +360,7 @@ const PROVIDER_SECTIONS: BYOKProviderSection[] = [
'cohere',
'xai',
'kimi',
'typesafe',
'fireworks',
'together',
'baseten',
Expand Down
225 changes: 225 additions & 0 deletions apps/sim/blocks/agent-evaluation.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,225 @@
/** @vitest-environment node */
import { resetEnvFlagsMock, setEnvFlags } from '@sim/testing'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import {
getEffectiveBlockOutputPaths,
getEffectiveBlockOutputs,
getEffectiveBlockOutputType,
} from '@/lib/workflows/blocks/block-outputs'
import { getBlockReferenceTags } from '@/lib/workflows/blocks/block-reference-tags'
import { evaluateSubBlockCondition } from '@/lib/workflows/subblocks/visibility'
import { AgentBlock } from '@/blocks/blocks/agent'
import { getAgentModelOptions, getModelOptions } from '@/blocks/utils'
import { getBaseModelProviders } from '@/providers/models'
import { Serializer } from '@/serializer'
import { useProvidersStore } from '@/stores/providers/store'
import type { BlockState } from '@/stores/workflows/workflow/types'

const { mockGetBlock } = vi.hoisted(() => ({ mockGetBlock: vi.fn() }))

vi.mock('@/blocks', () => ({ getBlock: mockGetBlock }))

describe('Agent evaluation configuration', () => {
afterEach(resetEnvFlagsMock)
beforeEach(() => {
mockGetBlock.mockReturnValue(AgentBlock)
})

it.each(['jev-1.13.0', 'jev-latest', 'jev-preview'])(
'shows native fields and credentials for %s',
(model) => {
const visible = AgentBlock.subBlocks
.filter((field) => evaluateSubBlockCondition(field.condition, { model }))
.map((field) => field.id)
expect(visible).toEqual(['model', 'apiKey', 'evaluationState', 'evaluationQuestions'])
}
)

it('keeps evaluation inputs configurable for a model reference', () => {
for (const field of AgentBlock.subBlocks.filter((field) => field.id.startsWith('evaluation'))) {
expect(evaluateSubBlockCondition(field.condition, { model: '<start.model>' })).toBe(true)
}
})

it.each([false, true])('shows TypeSafe credentials only when needed, hosted=%s', (hosted) => {
setEnvFlags({ isHosted: hosted })
const apiKey = AgentBlock.subBlocks.find((field) => field.id === 'apiKey')!
expect(evaluateSubBlockCondition(apiKey.condition, { model: 'jev-latest' })).toBe(!hosted)
})

it.each(['jev-1.13.0', '<start.model>', '{{MODEL_ID}}'])(
'exposes answers for %s in downstream selectors',
(model) => {
const values = { model: { value: model } }
expect(getEffectiveBlockOutputs('agent', values)).toHaveProperty('answers')
expect(getEffectiveBlockOutputPaths('agent', values)).toContain('answers')
expect(getEffectiveBlockOutputType('agent', 'answers', values)).toBe('json')
}
)

it('does not expose evaluation answers for a known chat model', () => {
expect(getEffectiveBlockOutputs('agent', { model: { value: 'gpt-4o' } })).not.toHaveProperty(
'answers'
)
})

it.each(['jev-1.13.0', '<start.model>'])(
'keeps answers accessible with a saved chat schema for %s',
(model) => {
const outputs = getEffectiveBlockOutputs('agent', {
model: { value: model },
responseFormat: {
value: { schema: { type: 'object', properties: { title: { type: 'string' } } } },
},
})
expect(outputs).toHaveProperty('answers')
if (model === 'jev-1.13.0') expect(outputs).not.toHaveProperty('title')
else expect(outputs).toHaveProperty('title')
}
)

it('shows Jev only in the model picker that supports evaluation inputs', () => {
useProvidersStore.getState().setProviderModels('base', Object.keys(getBaseModelProviders()))
expect(getAgentModelOptions().map((option) => option.id)).toContain('jev-1.13.0')
expect(getModelOptions().map((option) => option.id)).not.toContain('jev-1.13.0')
})

describe('evaluation answer references', () => {
const questions = {
category: { type: 'choice', instructions: 'Choose a category', criteria: { a: 'A', b: 'B' } },
rating: { type: 'score', instructions: 'Rate the result', criteria: ['Low', 'High'] },
passed: { type: 'noul', instructions: 'Did it pass?' },
}

it.each(['jev-1.13.0', 'jev-latest', 'jev-preview', '<start.model>', '{{MODEL_ID}}'])(
'exposes typed question fields for %s without an execution result',
(model) => {
const values = {
model: { value: model },
evaluationQuestions: { value: JSON.stringify(questions) },
responseFormat: {
value: { schema: { type: 'object', properties: { title: { type: 'string' } } } },
},
}
const tags = getBlockReferenceTags({
block: { id: 'agent-test', type: 'agent', name: 'Evaluate', subBlocks: values },
})
const fields = {
'answers.category.choice': 'string',
'answers.category.confidence': 'number',
'answers.category.probabilities': 'json',
'answers.category.type': 'string',
'answers.rating.score': 'number',
'answers.rating.confidence': 'number',
'answers.rating.legend': 'json',
'answers.passed.noul': 'number',
}
for (const [path, type] of Object.entries(fields)) {
expect(tags).toContain(`evaluate.${path}`)
expect(getEffectiveBlockOutputType('agent', path, values)).toBe(type)
}
expect(getEffectiveBlockOutputType('agent', 'answers', values)).toBe('json')
expect(getEffectiveBlockOutputType('agent', 'answers.category', values)).toBe('json')
expect(tags).not.toContain('evaluate.answers.passed.confidence')
expect(tags.includes('evaluate.title')).toBe(!model.startsWith('jev-'))
}
)

it.each([undefined, '', '{', '<start.questions>', '{{QUESTIONS}}', [], null, { unknown: {} }])(
'keeps the answers object selectable when questions cannot be inferred: %j',
(value) => {
const values = { model: { value: 'jev-latest' }, evaluationQuestions: { value } }
expect(getEffectiveBlockOutputPaths('agent', values)).toContain('answers')
expect(getEffectiveBlockOutputType('agent', 'answers', values)).toBe('json')
}
)

it('uses structured questions and follows edits without leaking fields into chat models', () => {
const values = {
model: { value: 'jev-latest' },
evaluationQuestions: { value: { result: questions.category } },
}
expect(getEffectiveBlockOutputPaths('agent', values)).toContain('answers.result.choice')
expect(
getEffectiveBlockOutputPaths('agent', {
...values,
evaluationQuestions: { value: { result: questions.passed } },
})
).not.toContain('answers.result.choice')
expect(
getEffectiveBlockOutputPaths('agent', {
...values,
model: { value: 'gpt-4o' },
}).some((path) => path.startsWith('answers'))
).toBe(false)
})

it('does not offer ambiguous reference paths for special question IDs', () => {
const values = {
model: { value: 'jev-latest' },
evaluationQuestions: {
value: {
'with.dot': questions.passed,
'with space': questions.passed,
'with[0]': questions.passed,
'<start.question>': questions.passed,
'valid-id_1': questions.passed,
},
},
}
const paths = getEffectiveBlockOutputPaths('agent', values)
expect(paths.filter((path) => path.startsWith('answers.'))).toEqual([
'answers.valid-id_1.noul',
'answers.valid-id_1.type',
])
expect(getEffectiveBlockOutputType('agent', 'answers', values)).toBe('json')
})

it.each(['type', 'properties', 'description', '__proto__'])(
'resolves the question named %s through schema properties',
(id) => {
const values = {
model: { value: 'jev-latest' },
evaluationQuestions: { value: { [id]: questions.passed } },
}
expect(getEffectiveBlockOutputPaths('agent', values)).toContain(`answers.${id}.noul`)
expect(getEffectiveBlockOutputType('agent', `answers.${id}.noul`, values)).toBe('number')
}
)
})

it.each([false, true])(
'serializes native fields without requiring messages, advanced=%s',
(advancedMode) => {
const values = {
model: 'jev-1.13.0',
apiKey: '{{TYPESAFE_API_KEY}}',
evaluationState: '42',
evaluationQuestions: '{"passed":{"type":"noul","instructions":"Did it pass?"}}',
messages: JSON.stringify([{ role: 'user', content: 'Old chat prompt' }]),
}
const block: BlockState = {
id: 'agent-test',
type: 'agent',
name: 'Evaluator',
position: { x: 0, y: 0 },
enabled: true,
advancedMode,
outputs: {},
subBlocks: Object.fromEntries(
Object.entries(values).map(([id, value]) => [
id,
{ id, value, type: AgentBlock.subBlocks.find((field) => field.id === id)!.type },
])
),
}
const result = new Serializer().serializeWorkflow({ [block.id]: block }, [], {}, {}, true)
expect(result.blocks[0].config.tool).toBe('typesafe')
expect(result.blocks[0].config.params).toMatchObject({
evaluationState: '42',
evaluationQuestions: values.evaluationQuestions,
})
expect(result.blocks[0].config.params).not.toHaveProperty('messages')
}
)
})
Loading
Loading