Working with Turbine Schema
Turbine Schema was formerly known as Turbine Extensible Data Schema, or TEDS.
Overview
Turbine Schema establishes a standardized data schema aimed at creating a unified framework to support seamless collaboration between security teams and tools. It promotes consistent data formats and schemas within Swimlane Turbine products, enhancing the detection, analysis, and response capabilities for security incidents. This standardization also simplifies the integration process for various cybersecurity tools, reducing complexity and minimizing integration effort.
This guide covers:
- What Turbine Schema is and why it matters
- How it relates to interfaces
- Where and how it is used across Turbine solutions
- Best practices, anti-patterns, and troubleshooting
Schema References
Field-level definitions for each solution area live in these documents:
- Turbine Schema Reference (Classic SOC)Turbine Schema Reference (Classic SOC) -- Alert, Email, Observable, Enrichment, and supporting objects for the SOC Solutions Bundle
- Turbine Schema Reference (AI SOC)Turbine Schema Reference (AI SOC) -- Extended Alert and Email objects for the AI SOC Solution
- Turbine Schema Reference (VRM)Turbine Schema Reference (VRM) -- Vulnerability Finding, Asset, and Remediation/Ticket objects for VRM
Interface Catalogs
Interfaces define the input/output contracts that use Turbine Schema objects. For interface documentation, see Working with InterfacesWorking with Interfaces.
How Turbine Schema Relates to Interfaces
Turbine Schema and interfaces serve complementary roles:
- Turbine Schema defines the data models -- the structure and fields of objects like Alert, Email, Observable, and Vulnerability Finding.
- Interfaces define the input/output contracts -- they specify which Turbine Schema objects a component accepts and produces.
When you apply an interface to a component, the interface configures the component's input and output schemas using Turbine Schema objects. This means components that use the same interface automatically share the same data format, making them interchangeable.
Where Turbine Schema Is Used
Turbine Schema schemas are integrated throughout Swimlane Turbine in several key areas:
SOC Solutions Bundle
The SOC Solutions Bundle uses Turbine Schema schemas extensively for security operations workflows:
- Alert Triage Solution: Processes alerts from SIEM, XDR, and EDR systems using the Alert object schema. Alerts are ingested via webhooks or API requests and transformed into standardized Alert objects following Turbine Schema conventions.
- Phishing Triage Solution: Processes suspected phishing emails using the Email and Phishing Email Report object schemas. Email data is extracted and structured according to Turbine Schema standards for consistent processing and analysis.
- Threat Intelligence Solution: Uses Observable and Enrichment object schemas to standardize threat intelligence data from various providers, ensuring consistent enrichment results across different sources.
AI SOC Solution
The AI SOC Solution extends the Classic SOC schemas with additional fields for:
- Enhanced alert context (priority, host criticality, MITRE D3FEND mappings, supporting evidence)
- Email authentication checks (SPF, DMARC, DKIM)
- Search-based ingestion parameters for alerts and emails
Vulnerability Response Management
The VRM Solution uses its own set of Turbine Schema objects for:
- Vulnerability Finding: Captures scan results from tools like Tenable, Qualys, and Rapid7 with CVSS/EPSS scoring, exploit intelligence, and remediation tracking
- Enriched Vulnerability Finding: Extends findings with asset criticality and zone context
- Asset: Represents managed assets in the inventory
- Remediation Item / Ticket: Tracks ITSM ticket creation and status
Application Field Naming
When building applications in Swimlane Turbine, you can follow Turbine Schema naming conventions for your field keys to ensure compatibility with Turbine Schema-based workflows and integrations. This is especially important when:
- Creating applications that will receive data from Turbine Schema-compliant sources
- Building custom solutions that integrate with any Solutions Bundle
- Ensuring data consistency across multiple systems and integrations
Playbook Actions
Turbine Schema schemas can be applied to playbook actions through Input Schema References. When configuring record actions (Create, Update/Create, Search), you can reference Turbine Schema-based schemas to ensure your playbooks accept and process data in standardized formats.
Business Use Case
For customers who prefer to develop their own solutions, data management can present a significant challenge. Standard data fields and naming conventions become essential for maintaining consistency and avoiding data loss or errors.
For instance, in scenarios where customers use a database like MongoDB Atlas Event Manager, even though Swimlane does not provide a direct solution for this specific database, the data still needs to be accurately saved to the database or application. A schema with consistent naming conventions is crucial; mismatches between field names can lead to lost or mishandled records. For customers unfamiliar with Turbine Schema references, it is important to ensure that their applications adhere to correct naming conventions to avoid any data errors and ensure reliable data management across systems.
How to Use Turbine Schema in Your Workflows
Using Turbine Schema in Applications
When building applications that will work with Turbine Schema-based data:
- Follow Turbine Schema Naming Conventions: Use the field keys defined in the relevant schema reference document when creating application fields. For example, if creating an Alert application, use alert_uid as the field key for the unique identifier, not alertId or alert-id.
- Match Field Types: Ensure your application field types match the Turbine Schema types. For example:
- Use String fields for alert_title, alert_description
- Use Date & Time fields for alert_created_timestamp, alert_start_timestamp
- Use Multi-Select or Array fields for alert_categories, alert_impacted_hostnames
- Use Reference fields for nested objects like observables or alert_rules
- Required vs Optional Fields: Mark fields as required based on the Turbine Schema requirements. Fields marked as "Required" should be required in your application, while "Recommended" and "Optional" fields can be optional.
Using Turbine Schema in Playbooks
When building playbooks that process Turbine Schema-formatted data:
- Apply Interfaces: Use interfaces that implement Turbine Schema schemas. These interfaces automatically configure your component's input and output schemas to match Turbine Schema standards.
- Reference Schemas: When configuring record actions, you can reference Turbine Schema-based input schemas to ensure your playbooks accept data in the correct format.
- Data Transformation: Use transformation functions to map incoming data to Turbine Schema field names if your source data uses different naming conventions.
Example Workflow
Here is an example of how Turbine Schema is used in an Alert Triage workflow:
- Alert Ingestion: A webhook receives an alert from a SIEM system
- Schema Application: The alert data is validated against the Alert Turbine Schema
- Observable Extraction: Observables (IPs, URLs, file hashes) are extracted and structured using the Observable schema
- Enrichment: Each observable is enriched using threat intelligence providers, with results following the Enrichment schema
- Case Creation: A case is created in the Case and Incident Management application using Turbine Schema field names
- Analysis: The case is analyzed using Hero AI or manual review, with all data following Turbine Schema conventions
This standardized approach ensures that data flows seamlessly between different components and systems, regardless of the underlying technology or vendor.
Guidelines for Attribute Names
- Attribute names must be valid UTF-8 sequences.
- Use lowercase for all attribute names.
- Separate words with underscores.
- Apply present tense unless the attribute refers to historical information.
- Use singular or plural forms appropriately to match the field content.
- Example: Use events_per_sec instead of event_per_sec.
- If an attribute represents multiple entities, use a pluralized name and set the value type as an array.
- Example: process.loaded_modules stores a list of module names.
- Avoid word repetition.
- Example: Instead of host.host_ip, use host.ip.
- Minimize abbreviations, with exceptions for commonly recognized terms (for example, ip, os, geo).
Attribute Levels
The schema categorizes attributes into three levels:
- Core Attributes: Common across all use cases, designated as either Required or Recommended.
- Optional Attributes: Relevant to specific use cases or allow flexibility based on the context.
- Reserved Attributes: Managed by the logging system and should not be used in event data.
Extending the Schema
The Open Cybersecurity Schema Framework allows for extensions through additional attributes, objects, and event classes.
To extend the schema:
- Create a new directory mirroring the top-level schema directory structure.
- This directory can include the following files and subdirectories:
- categories.json: Defines a new event category and reserves a range of class IDs.
- dictionary.json: Defines new attributes.
- events/: Contains definitions for new event classes.
- objects/: Holds definitions for new objects.
Best Practices
Field Naming Consistency
- Always use lowercase: Field keys must be lowercase (for example, alert_uid, not Alert_UID or alertUid)
- Use underscores: Separate words with underscores (for example, alert_created_timestamp, not alertCreatedTimestamp or alert-created-timestamp)
- Follow Turbine Schema conventions: Use the exact field keys defined in the reference documents to ensure compatibility
- Avoid abbreviations: Unless they are commonly recognized (for example, ip, os, geo), spell out full words
Schema Compliance
- Validate against schemas: When building custom solutions, validate your data structures against Turbine Schema before deployment
- Handle optional fields: Design your workflows to handle missing optional fields gracefully
- Required fields: Always include required fields; missing required fields will cause validation errors
- Type matching: Ensure data types match the schema definitions (for example, arrays for multi-value fields, datetime for timestamps)
Integration Tips
- Start with Solutions Bundles: If you are new to Turbine Schema, start by using a Solutions Bundle, which already implements Turbine Schema correctly
- Use Interfaces: Apply interfaces to your components to automatically get Turbine Schema-compliant schemas
- Test transformations: When mapping data from external sources to Turbine Schema format, test your transformations thoroughly
- Document deviations: If you need to extend Turbine Schema, document your extensions clearly
Common Anti-Patterns
Understanding what not to do is just as important as following best practices. The following examples show common mistakes and how to correct them.
Incorrect Field Naming
Problem: Using camelCase or mixed case instead of snake_case.
Wrong:
{
"alertUid": "alert-12345",
"alertCreatedTimestamp": "2025-01-15T10:30:00Z",
"alertImpactedHostnames": ["host1", "host2"]
}Correct:
{
"alert_uid": "alert-12345",
"alert_created_timestamp": "2025-01-15T10:30:00Z",
"alert_impacted_hostnames": ["host1", "host2"]
}Why it fails: Field names are case-sensitive. Components expecting alert_uid will not find alertUid, causing data to be lost or validation errors.
Wrong Field Types
Problem: Using strings for array fields or incorrect data types.
Wrong:
{
"alert_categories": "Phishing, Malware",
"alert_impacted_hostnames": "WORKSTATION-01",
"alert_risk_score": "85"
}Correct:
{
"alert_categories": ["Phishing", "Malware"],
"alert_impacted_hostnames": ["WORKSTATION-01"],
"alert_risk_score": 85
}Why it fails: Array fields must be arrays, not comma-separated strings. Integer fields must be numbers, not strings. Type mismatches cause validation errors and prevent proper data processing.
Missing Required Fields
Problem: Omitting required fields from Turbine Schema objects.
Wrong:
{
"alert_title": "Suspicious Activity",
"alert_severity": "High"
}Correct:
{
"alert_uid": "alert-12345",
"alert_title": "Suspicious Activity",
"alert_severity": "High",
"raw_alert": {}
}Why it fails: Required fields (alert_uid and raw_alert for Alert objects) must always be present. Missing required fields cause schema validation failures and prevent data ingestion.
Incorrect Nested Object Structure
Problem: Not following the correct structure for nested objects like observables or enrichments.
Wrong:
{
"observables": {
"observable_type": "ipv4_public",
"observable_value": "203.0.113.1"
}
}Correct:
{
"observables": [
{
"observable_type": "ipv4_public",
"observable_value": "203.0.113.1"
}
]
}Why it fails: observables must be an array, even for a single observable. Using an object instead of an array causes type validation errors.
Incorrect Observable Type Values
Problem: Using invalid values for observable_type field.
Wrong:
{
"observable_type": "IP",
"observable_value": "203.0.113.1"
}Correct:
{
"observable_type": "ipv4_public",
"observable_value": "203.0.113.1"
}Why it fails: observable_type must use exact Turbine Schema values: ipv4_public, ipv4_private, ipv6_public, ipv6_private, url, domain, email, sha256, sha1, md5, or file. Invalid types cause validation errors and prevent observable processing.
Date Format Inconsistencies
Problem: Using incorrect date/time formats.
Wrong:
{
"alert_created_timestamp": "2025-01-15 10:30:00",
"alert_start_timestamp": "Jan 15, 2025 10:30 AM"
}Correct:
{
"alert_created_timestamp": "2025-01-15T10:30:00Z",
"alert_start_timestamp": "2025-01-15T10:28:15Z"
}Why it fails: All datetime fields must use ISO 8601 format with UTC timezone (Z suffix). Incorrect formats cause parsing errors and time-based queries to fail.
Mixing Naming Conventions
Problem: Inconsistent naming within the same object.
Wrong:
{
"alert_uid": "alert-12345",
"alertTitle": "Suspicious Activity",
"alert-created-timestamp": "2025-01-15T10:30:00Z"
}Correct:
{
"alert_uid": "alert-12345",
"alert_title": "Suspicious Activity",
"alert_created_timestamp": "2025-01-15T10:30:00Z"
}Why it fails: All fields must consistently use snake_case. Mixing conventions causes some fields to be unrecognized and data to be lost.
Troubleshooting Common Issues
Data Not Appearing in Applications
Problem: Data ingested via Turbine Schema is not appearing in your application records.
Solutions:
- Verify field keys match Turbine Schema conventions exactly (case-sensitive, underscore-separated)
- Check that field types match the schema (for example, array fields for multi-value data)
- Ensure required fields are present in your data
- Validate your data structure against the Turbine Schema before ingestion
Schema Validation Errors
Problem: Playbook actions fail with schema validation errors.
Solutions:
- Review the error message to identify which field is causing the issue
- Compare your data structure to the Turbine Schema definition
- Check for typos in field names (for example, alert_uid vs alert_ui)
- Ensure nested objects follow the correct structure (for example, Observable objects within arrays)
Integration Compatibility Issues
Problem: Components or connectors are not working together as expected.
Solutions:
- Verify all components use the same interface version
- Check that input/output schemas match between connected components
- Review the interface documentation to ensure you are using compatible interfaces
- Test components individually before integrating them into larger workflows
Observable Enrichment Failures
Problem: Observable enrichments are not being applied or are missing from results.
Solutions:
- Verify observable_type uses valid Turbine Schema values (for example, ipv4_public, not IP or ip)
- Ensure observable_value is properly formatted (for example, valid IP address, URL, or hash)
- Check that enrichment provider components are correctly configured
- Verify enrichment results follow the Enrichment schema structure:
{
"enrichment_type": "reputation",
"enrichment_provider": "VirusTotal",
"enrichment_verdict": "Malicious",
"enrichment_timestamp": "2025-01-15T10:00:00Z"
}- Review playbook execution logs to identify which enrichment step failed
Date/Time Format Issues
Problem: Timestamp fields are not being recognized or parsed correctly.
Solutions:
- Ensure all datetime fields use ISO 8601 format: YYYY-MM-DDTHH:mm:ssZ
- Verify timezone is specified (use Z for UTC or +HH:mm for other timezones)
- Check that datetime fields are strings, not Date objects in JSON
- For relative times (for example, "4 hours ago"), ensure your transformation logic converts them to absolute timestamps before storing
Array vs Single Value Confusion
Problem: Data appears as a single value when it should be an array, or vice versa.
Solutions:
- Always use arrays for multi-value fields, even when there is only one value
- Never use comma-separated strings for array fields
- Check field definitions in your application to ensure array fields are configured as Multi-Select or Array types
- Use transformation functions to convert comma-separated strings to arrays if needed
Nested Object Structure Problems
Problem: Nested objects like observables, enrichments, or detection rules are not structured correctly.
Solutions:
- Verify nested arrays contain objects, not primitives
- Ensure required fields are present in nested objects (for example, observable_type and observable_value for Observable objects)
- Check that nested object structures match the Turbine Schema exactly (field names, types, nesting levels)
- Use the Solutions Bundle interfaces as reference implementations
Field Name Typos and Case Sensitivity
Problem: Fields are not being recognized due to typos or case mismatches.
Solutions:
- Double-check field names against the relevant Turbine Schema Reference document (field names are case-sensitive)
- Common typos to avoid:
- alert_uid vs alert_ui (missing "d")
- observable_type vs observableType (wrong case)
- email_from_address vs email_from (missing "_address")
- Use IDE autocomplete or schema validation tools to catch typos early
- Compare your field names character-by-character with the schema definitions
Missing Raw Data Fields
Problem: The raw_alert or raw_email field is missing or incorrectly formatted.
Solutions:
- Always include raw_alert (for Alert objects) or raw_email (for Email objects) as required fields
- Store the complete original payload in the raw field
- Ensure the raw field is a JSON object, not a string (unless your system requires string serialization)
- Preserve all original data to enable forensic analysis and debugging