Langfuse AI employee

A Superkind AI employee that checks traces, explains failed generations, analyses scores, and documents incidents in Langfuse. Message it in Teams or Outlook, it works in Langfuse and reports back with the result. Sensitive actions wait for your approval.

  • Hosted in the EU
  • GDPR data processing agreement
  • Ready to start today
  • Works in Teams, Slack and more

What is a Langfuse AI employee?

A Langfuse AI employee connects to your Langfuse account and completes work in it. It checks traces and observations, explains failed generations, analyses scores, and documents incidents. Superkind builds it with your company knowledge, quality rules, and technical context.

Unlike Zapier or Make, there is nothing to configure. You describe the outcome in plain English, the AI employee picks the right Langfuse actions, connects them with Teams, Jira, or Linear, and asks for your approval before sensitive changes.

About Langfuse

Langfuse is a platform for LLM observability, prompt management, evaluations, datasets, and experiments.

  1. You ask in Teams

    Describe in plain English which error or quality problem you want to investigate.

  2. Superkind picks the actions

    Selects the right Langfuse actions and connects them with your other systems.

  3. Langfuse

    Superkind works in Langfuse

    Checks traces, observations, scores, and prompts using your real data.

  4. Superkind reports back

    Delivers the cause, evidence, and next steps in Teams or Outlook.

Try asking

What can you ask Superkind to do in Langfuse?

Simply write what you want checked or completed, and Superkind takes it from there.

Check today’s deployment in Langfuse and show me every failed trace with their common cause.

you, to @Superkind

Use the observations to explain why checkout_agent has been timing out more often since this morning.

you, to @Superkind

Create a Jira ticket for this error with affected generations, release, and your root cause analysis.

you, to @Superkind

Let me know here as soon as a new incident appears in Langfuse, with trace, score, and likely cause.

you, to @Superkind
How it works

How does Superkind work with Langfuse?

  1. Native integrations and connectors for 1,000+ tools

    1Connect your systems

    Connect Langfuse with the required API keys. Add Teams and the systems your team already uses. Superkind configures access with minimum permissions. Your AI employee is then available in Teams.

  2. Lena Hoffmann9:12 AM

    @Superkind check the failed traces in today’s release, explain the cause, and create a score for the incident.

    @SuperkindApp9:13 AM

    On it. I am checking the traces and observations in Langfuse and creating the score.

    2Tell Superkind what you need

    Message it in Teams like a colleague and say which traces or scores you want checked. Superkind understands your context and picks the right Langfuse actions. It can connect results with Jira or Linear. You do not need to build a workflow.

  3. @SuperkindApp9:14 AM
    • 12 release traces checked
    • Error: checkout_agent timeout
    • Score latency_regression created
    Langfuse
    Score latency_regression12 traces · critical
    Waiting for your approvalMove prompt label to production?

    3Superkind operates, you approve

    Superkind works in Langfuse and reports the cause, evidence, and next steps back in Teams. Read only checks run independently. Changes to prompts, datasets, or access wait for your approval. Every step is logged.

Actions

What can Superkind do in Langfuse?

Ask in plain English from Teams or Outlook. Superkind picks the right Langfuse actions, runs the work, and reports back. No workflows to build.

  • Traces and observations

    Search observations

    Finds spans, generations, and events by time range, name, level, environment, or trace ID.

  • Traces and observations

    Reconstruct trace

    Combines the observations for a trace ID into a complete execution path.

  • Traces and observations

    Check generations

    Checks generations for errors, latency, model, token usage, and cost.

  • Traces and observations

    Analyse sessions

    Groups traces by session ID and explains recurring problems across a conversation.

  • Traces and observations

    Explain error cause

    Compares failed observations and identifies shared status messages, models, or releases.

  • Metrics

    Aggregate costs

    Aggregates Langfuse costs by model, trace name, release, or time range.

  • Metrics

    Compare latency

    Compares latency and time to first token across models, releases, or environments.

  • Metrics

    Analyse token usage

    Shows input, output, and total token usage for selected Langfuse dimensions.

  • Metrics

    Measure trace volume

    Measures trace volume by application, user, model, or time range.

  • Scores

    Search scores

    Finds scores by name, data type, source, value, environment, or time range.

  • Scores

    Explain score trends

    Explains changes in quality values across releases, models, or prompt versions.

  • ScoresNeeds approval

    Create score

    Creates a numeric, boolean, categorical, or text score on a trace or observation.

  • Scores

    Check score configs

    Lists score configs with data type, categories, and allowed values.

  • Prompts

    List prompts

    Lists Langfuse prompts with type, version, labels, and tags.

  • Prompts

    Get prompt version

    Fetches a specific prompt version or the version behind a label.

  • Prompts

    Compare prompt impact

    Compares scores, cost, and latency across linked prompt versions.

  • PromptsNeeds approval

    Create prompt version

    Creates a new Langfuse prompt version with config, labels, and tags.

  • Datasets and experiments

    List datasets

    Lists datasets with description, metadata, and creation date.

  • Datasets and experiments

    Check dataset items

    Reads inputs, expected outputs, metadata, and status of individual dataset items.

  • Datasets and experimentsNeeds approval

    Create dataset

    Creates a Langfuse dataset for repeatable evaluations.

  • Datasets and experimentsNeeds approval

    Create dataset item

    Adds an item with input, expected output, and metadata to a dataset.

  • Datasets and experiments

    Compare experiments

    Compares experiment runs using their items, scores, cost, and latency.

  • Datasets and experiments

    Find experiment failures

    Finds experiment items with divergent outputs or low scores.

  • Other

    Check project access

    Checks Langfuse project roles and permissions without changing access.

  • OtherNeeds approval

    Invite member

    Invites a member to the Langfuse organisation with a defined role.

  • OtherNeeds approval

    Delete project

    Deletes a Langfuse project with its traces, scores, prompts, and datasets.

Works with your stack

Superkind uses Langfuse together with your other systems

Ask for outcomes across tools. The AI employee checks Langfuse, connects technical context from GitHub, Jira, or Linear, and reports the result in Teams, Outlook, or Slack.

Matching AI employees

Superkind AI employees that work with Langfuse

Every role brings its expertise and uses Langfuse as one of its tools.

Companies working with Superkind

CG Group
CG Real Estate
Ecobuilding
Nivocare
Lindenstrom
Tylrus
Lorvan
Voelpker
Voelpker
Lorvan
FAQ

Frequently asked questions

Everything you need to know about your AI employee for Langfuse.

Yes. Superkind connects to your Langfuse project through a managed connector. Your AI employee can then check traces, observations, metrics, scores, prompts, datasets, and experiments from Teams or Outlook. It receives only the agreed permissions and reports results back with the relevant evidence from Langfuse.

An admin provides the Langfuse API credentials for the selected project. Superkind configures the connector, tests the connection, and limits permissions to the agreed tasks. You can then reach your AI employee in Teams or Outlook. We test a real query together before your team starts using the connection.

It can search observations, reconstruct traces, explain failed generations, and analyse metrics for cost, tokens, or latency. It can also check scores, compare prompt versions, read datasets, and summarise experiment runs. Write actions such as creating scores or dataset items only run according to the approval rules you define.

No. With Zapier or Make, you build and maintain triggers, filters, and individual steps. With Superkind, you simply describe the desired outcome in Teams or Outlook. The AI employee picks the right Langfuse actions itself, connects them with GitHub, Jira, or Linear when needed, and asks when information is missing.

Only with your approval. Reading traces, comparing scores, and explaining errors can run independently. New prompt versions, dataset items, members, or deleting a project wait for an explicit yes from a responsible person in Teams. Your team decides together with Superkind which Langfuse actions count as sensitive.

Superkind is hosted in the EU and signs a GDPR data processing agreement with you. The connector uses minimum Langfuse permissions for the agreed tasks. Your data is not used for model training. Access and completed actions are logged so you can trace which traces, scores, prompts, or datasets the AI employee processed.

Langfuse offers a free starting tier. Paid cloud plans begin in roughly the low double digits per month, while broader plans sit in the low hundreds of euros per month. Zapier starts at about 20 euros and Make at about 10 euros per month, both plus setup time. Superkind is priced per use case, a fraction of a full-time hire.

Putting your AI to workContact ustogether