What a JSON application line looks like
The JSON sample below is fed verbatim into the engine to produce every parser on this page.
{"ts":"2026-07-03T14:22:15.003Z","level":"error","service":"checkout","msg":"payment failed","order_id":"ord_9f3c","user":"jdoe","ip":"203.0.113.45","gateway":{"name":"stripe","code":"card_declined"}}
{"ts":"2026-07-03T14:22:18.220Z","level":"info","service":"auth","msg":"login ok","user":"berkay","ip":"192.0.2.10","mfa":true} Detected fields
The engine classified this sample as json and consolidated 10 fields across 2 lines. Fields marked literal were identical on every sample line, so they are baked into the pattern as anchors rather than captured.
- ts : timestamp
- level : severity
- service : quoted_string
- msg : quoted_string
- order_id : quoted_string
- user : username
- ip : ipv4
- gateway_name : quoted_string
- gateway_code : quoted_string
- mfa : literal
Regex (named capture groups)
# sample: {"ts":"2026-07-03T14:22:15.003Z","level":"error","service":"checkout","msg":"payment failed","order_id":"ord_9f3c","user":"jdoe","ip":"203.0.113.45","gateway":{"name":"stripe","code":"card_declined"}}
# groups: ts=2026-07-03T14:22:15.003Z, level=error, service=checkout, msg=payment failed, order_id=ord_9f3c, user=jdoe, ip=203.0.113.45, gateway_name=stripe, gateway_code=card_declined
^(?=.*?"ts":"(?<ts>[^"]*)")(?=.*?"level":"(?<level>[^"]*)")(?=.*?"service":"(?<service>[^"]*)")(?=.*?"msg":"(?<msg>[^"]*)")(?=.*?"order_id":"(?<order_id>[^"]*)"|)(?=.*?"user":"(?<user>[^"]*)")(?=.*?"ip":"(?<ip>[^"]*)")(?=.*?"name":"(?<gateway_name>[^"]*)"|)(?=.*?"code":"(?<gateway_code>[^"]*)"|)(?=.*?"mfa":(?<mfa>[A-Za-z]+)|).*$ - note input is JSON — use a JSON parser (jq, Logstash json filter, …) instead of a regex where possible
- note a single linear template could not reproduce every input line — fields are captured with order-independent lookaheads instead
Grok pattern (Logstash / Elastic)
# custom patterns
JSON_NOTDQUOTE [^"]*
\{"ts":"%{TIMESTAMP_ISO8601:ts}","level":"%{LOGLEVEL:level}","service":"%{JSON_NOTDQUOTE:service}","msg":"%{JSON_NOTDQUOTE:msg}(?:","order_id":"%{JSON_NOTDQUOTE:order_id})?","user":"%{USERNAME:user}","ip":"%{IPV4:ip}(?:","gateway":\{"name":"%{JSON_NOTDQUOTE:gateway_name})?(?:","code":"%{JSON_NOTDQUOTE:gateway_code})?(?:","mfa":%{GREEDYDATA:mfa})? - note json input — consider the Logstash json codec/filter instead of grok
- note 4 optional field(s) wrapped in (?:…)? inline regex — grok has no native optional syntax
- note custom patterns emitted — save the '# custom patterns' block to a file in your patterns_dir
Wazuh decoder (OS_Regex XML)
<!--
Generated by LogForge - Wazuh decoder (OS_Regex dialect, not PCRE)
sample: {"ts":"2026-07-03T14:22:15.003Z","level":"error","service":"checkout","msg":"payment failed","order_id":"ord_9f3c","user":"jdoe","ip":"203.0.113.45","gateway":{
test with: /var/ossec/bin/wazuh-logtest
-->
<decoder name="json-json">
<prematch>^{</prematch>
<plugin_decoder>JSON_Decoder</plugin_decoder>
</decoder>
<!-- ============================================================
ALERT RULE (starter) — put this in a RULES file, e.g.
/var/ossec/etc/rules/local_rules.xml. Decoders and rules live
in SEPARATE files. The rule matches the decoder above through
<decoded_as>; set <level> and add <field>/<match> conditions so
it alerts only on the events you care about. Rule ids 100000+
are the user range — change them if they collide with yours.
============================================================ -->
<group name="json,">
<rule id="100000" level="3">
<decoded_as>json-json</decoded_as>
<description>json: decoded event</description>
</rule>
</group>
- note JSON input: emitted a JSON_Decoder plugin decoder — Wazuh extracts every key automatically as dynamic fields (nested keys become dotted names)
- note field "ip" mapped to Wazuh conventional field "srcip"
- note field names above are what the other LogForge generators use; JSON_Decoder will use the raw JSON keys instead
- note added a starter alert <rule> (level 3, matched to the decoder via <decoded_as>) — put it in a RULES file (not the decoders file), set the level, and add <field>/<match> conditions; the commented example child rule shows the pattern
- note decoder order and prematch specificity may need site-specific tuning (other decoders in your ruleset can shadow these) — validate with /var/ossec/bin/wazuh-logtest
Wazuh's OS_Regex is not PCRE — a bare . is a literal dot and \. matches any character.
Test Wazuh OS_Regex patterns →
rsyslog template / liblognorm rulebase
version=2
# json — liblognorm v2 rulebase (generated by LogForge)
# Usage with rsyslog (mmnormalize runs liblognorm):
# module(load="mmnormalize")
# action(type="mmnormalize" rulebase="/etc/rsyslog.d/json.rb" useRawMsg="on")
# Literal "%" is escaped as "%%"; raw tabs are written as \x09.
rule=json:{"ts":"%ts:date-rfc5424%","level":"%level:char-to{"extradata":"\""}%","service":"%service:char-to{"extradata":"\""}%","msg":"%msg:char-to{"extradata":"\""}%","order_id":"%order_id:char-to{"extradata":"\""}%","user":"%user:char-to{"extradata":"\""}%","ip":"%ip:ipv4%","gateway":{"name":"%gateway_name:char-to{"extradata":"\""}%","code":"%gateway_code:char-to{"extradata":"\""}%","mfa":%mfa:char-to{"extradata":"\""}%"}}
rule=json:{"ts":"%ts:date-rfc5424%","level":"%level:char-to{"extradata":"\""}%","service":"%service:char-to{"extradata":"\""}%","msg":"%msg:char-to{"extradata":"\""}%","user":"%user:char-to{"extradata":"\""}%","ip":"%ip:ipv4%"
- note json structure: rsyslog mmjsonparse handles CEE/JSON natively — consider action(type="mmjsonparse") instead of this rulebase
- note trailing literal "\"}}" reconstructed from line 1
- note chosen parser types: ts=date-rfc5424, level=char-to("), service=char-to("), msg=char-to("), order_id=char-to("), user=char-to("), ip=ipv4, gateway_name=char-to("), gateway_code=char-to("), mfa=char-to(")
- note optional columns (order_id, gateway_name, gateway_code, mfa): liblognorm has no optional parts within a single rule — emitted a second rule variant with only the always-present columns (max 2 variants; lines with other column combinations will not match and need extra rule= lines)
Splunk
# props.conf (search-time extraction)
[<REPLACE_WITH_SOURCETYPE>]
EXTRACT-logforge = (?=.*?"ts":"(?<ts>[^"]*)")(?=.*?"level":"(?<level>[^"]*)")(?=.*?"service":"(?<service>[^"]*)")(?=.*?"msg":"(?<msg>[^"]*)")(?=.*?"order_id":"(?<order_id>[^"]*)"|)(?=.*?"user":"(?<user>[^"]*)")(?=.*?"ip":"(?<ip>[^"]*)")(?=.*?"name":"(?<gateway_name>[^"]*)"|)(?=.*?"code":"(?<gateway_code>[^"]*)"|)(?=.*?"mfa":(?<mfa>[A-Za-z]+)|).*
# Quick search-time test in SPL:
# | rex field=_raw "(?=.*?\"ts\":\"(?<ts>[^\"]*)\")(?=.*?\"level\":\"(?<level>[^\"]*)\")(?=.*?\"service\":\"(?<service>[^\"]*)\")(?=.*?\"msg\":\"(?<msg>[^\"]*)\")(?=.*?\"order_id\":\"(?<order_id>[^\"]*)\"|)(?=.*?\"user\":\"(?<user>[^\"]*)\")(?=.*?\"ip\":\"(?<ip>[^\"]*)\")(?=.*?\"name\":\"(?<gateway_name>[^\"]*)\"|)(?=.*?\"code\":\"(?<gateway_code>[^\"]*)\"|)(?=.*?\"mfa\":(?<mfa>[A-Za-z]+)|).*" - note regex: input is JSON — use a JSON parser (jq, Logstash json filter, …) instead of a regex where possible
- note regex: a single linear template could not reproduce every input line — fields are captured with order-independent lookaheads instead
- note EXTRACT-<class> names must be unique within a sourcetype stanza — rename EXTRACT-logforge if you already use that class for this sourcetype
- note a timestamp field was detected: this EXTRACT only makes it a searchable field. To set the event _time at index time, add TIME_PREFIX and TIME_FORMAT to this props.conf stanza (TIME_FORMAT uses Splunk strptime, e.g. %Y-%m-%dT%H:%M:%S) — this generator does not guess the strptime format.
ES ingest
PUT _ingest/pipeline/json
{
"description": "LogForge-generated ingest pipeline for json",
"processors": [
{
"json": {
"field": "message",
"add_to_root": true
}
}
]
} - note grok: json input — consider the Logstash json codec/filter instead of grok
- note grok: 4 optional field(s) wrapped in (?:…)? inline regex — grok has no native optional syntax
- note grok: custom patterns emitted — save the '# custom patterns' block to a file in your patterns_dir
- note json structure: emitted a { json: { field: "message", add_to_root: true } } processor — it parses the JSON line and lifts every field to the top level of the record, so a grok pattern is unnecessary
- note test in Kibana Dev Tools with: POST _ingest/pipeline/json/_simulate (supply a docs[] array whose _source.message holds a sample line)
Graylog
# Grok patterns to add under System > Grok Patterns:
# (Graylog needs these custom patterns installed globally BEFORE the rule/extractor below will work.)
# JSON_NOTDQUOTE [^"]*
# --- Graylog processing pipeline rule (primary) ---
# Paste under System > Pipelines > Manage rules, then attach the rule to a pipeline stage.
rule "json-parse"
when
has_field("message")
then
let gp = grok(pattern: "\\{\"ts\":\"%{TIMESTAMP_ISO8601:ts}\",\"level\":\"%{LOGLEVEL:level}\",\"service\":\"%{JSON_NOTDQUOTE:service}\",\"msg\":\"%{JSON_NOTDQUOTE:msg}(?:\",\"order_id\":\"%{JSON_NOTDQUOTE:order_id})?\",\"user\":\"%{USERNAME:user}\",\"ip\":\"%{IPV4:ip}(?:\",\"gateway\":\\{\"name\":\"%{JSON_NOTDQUOTE:gateway_name})?(?:\",\"code\":\"%{JSON_NOTDQUOTE:gateway_code})?(?:\",\"mfa\":%{GREEDYDATA:mfa})?", value: to_string($message.message), only_named_captures: true);
set_fields(gp);
end
# --- Graylog import-ready extractor JSON (secondary) ---
# Save as a .json file and import under System > Inputs > (input) > Manage extractors > Actions > Import extractors.
{
"extractors": [
{
"title": "json",
"extractor_type": "grok",
"converters": [],
"order": 0,
"cursor_strategy": "copy",
"source_field": "message",
"target_field": "",
"extractor_config": {
"grok_pattern": "\\{\"ts\":\"%{TIMESTAMP_ISO8601:ts}\",\"level\":\"%{LOGLEVEL:level}\",\"service\":\"%{JSON_NOTDQUOTE:service}\",\"msg\":\"%{JSON_NOTDQUOTE:msg}(?:\",\"order_id\":\"%{JSON_NOTDQUOTE:order_id})?\",\"user\":\"%{USERNAME:user}\",\"ip\":\"%{IPV4:ip}(?:\",\"gateway\":\\{\"name\":\"%{JSON_NOTDQUOTE:gateway_name})?(?:\",\"code\":\"%{JSON_NOTDQUOTE:gateway_code})?(?:\",\"mfa\":%{GREEDYDATA:mfa})?",
"named_captures_only": true
},
"condition_type": "none",
"condition_value": ""
}
],
"version": "5.0.0"
} - note grok: json input — consider the Logstash json codec/filter instead of grok
- note grok: 4 optional field(s) wrapped in (?:…)? inline regex — grok has no native optional syntax
- note grok: custom patterns emitted — save the '# custom patterns' block to a file in your patterns_dir
- note 1 custom grok pattern(s) (JSON_NOTDQUOTE) must be installed globally first under System > Grok Patterns — see the block at the top of the output
- note primary artifact is the processing-pipeline rule; the extractor JSON is an equivalent import-ready alternative for the classic extractor UI
Datadog
logforge_rule \{"ts":"%{date("yyyy-MM-dd'T'HH:mm:ss.SSSZ"):ts}","level":"%{word:level}","service":"%{quotedString:service}","msg":"%{quotedString:msg}(?:","order_id":"%{quotedString:order_id})?","user":"%{notSpace:user}","ip":"%{ipv4:ip}(?:","gateway":\{"name":"%{quotedString:gateway_name})?(?:","code":"%{quotedString:gateway_code})?(?:","mfa":%{notSpace:mfa})? - note emitted rule name is "logforge_rule"; rename it to match your "json" convention if desired
- note json input — Datadog can parse JSON logs automatically; a Grok Parser is only needed for non-JSON message bodies
- note 4 optional field(s) wrapped in (?:…)? — Datadog Grok has no native optional matcher; a chain of optional columns may need a Helper Rule per shape
- note paste this line into a Grok Parser processor in a Datadog Log Pipeline; matchers are anchored left-to-right and rule whitespace matches log whitespace. Complex or multi-shape logs may need Helper Rules.
Fluent Bit
[PARSER]
Name json
Format regex
Regex ^(?=.*?"ts":"(?<ts>[^"]*)")(?=.*?"level":"(?<level>[^"]*)")(?=.*?"service":"(?<service>[^"]*)")(?=.*?"msg":"(?<msg>[^"]*)")(?=.*?"order_id":"(?<order_id>[^"]*)"|)(?=.*?"user":"(?<user>[^"]*)")(?=.*?"ip":"(?<ip>[^"]*)")(?=.*?"name":"(?<gateway_name>[^"]*)"|)(?=.*?"code":"(?<gateway_code>[^"]*)"|)(?=.*?"mfa":(?<mfa>[A-Za-z]+)|).*$
Time_Key ts
Time_Format %Y-%m-%dT%H:%M:%S.%LZ
# Fluentd <parse> block:
# <parse>
# @type regexp
# expression /^(?=.*?"ts":"(?<ts>[^"]*)")(?=.*?"level":"(?<level>[^"]*)")(?=.*?"service":"(?<service>[^"]*)")(?=.*?"msg":"(?<msg>[^"]*)")(?=.*?"order_id":"(?<order_id>[^"]*)"|)(?=.*?"user":"(?<user>[^"]*)")(?=.*?"ip":"(?<ip>[^"]*)")(?=.*?"name":"(?<gateway_name>[^"]*)"|)(?=.*?"code":"(?<gateway_code>[^"]*)"|)(?=.*?"mfa":(?<mfa>[A-Za-z]+)|).*$/
# time_key ts
# time_format %Y-%m-%dT%H:%M:%S.%LZ
# </parse>
- note regex: input is JSON — use a JSON parser (jq, Logstash json filter, …) instead of a regex where possible
- note regex: a single linear template could not reproduce every input line — fields are captured with order-independent lookaheads instead
- note Time_Key set to "ts"; Time_Format "%Y-%m-%dT%H:%M:%S.%LZ" is a best-effort strptime derived from the sample shape — verify it against your data (Fluent Bit uses %L for fractional seconds and %z for numeric offsets)
Vector
[transforms.json_parse]
type = "remap"
inputs = ["REPLACE_WITH_SOURCE"]
source = '''
. = parse_json!(.message)
''' - note the reused regex needs lookahead, which the Rust regex crate behind parse_regex rejects; used parse_json! on .message (also the idiomatic Vector parser for JSON logs)
Loki
# promtail pipeline for "json" (generated by LogForge)
# Add these stages under a scrape_config in your promtail config:
# scrape_configs:
# - job_name: json
# pipeline_stages:
# (the stages below are indented to sit under pipeline_stages)
pipeline_stages:
- json:
expressions:
ts: 'ts'
level: 'level'
service: 'service'
msg: 'msg'
order_id: 'order_id'
user: 'user'
ip: 'ip'
gateway_name: 'gateway.name'
gateway_code: 'gateway.code'
mfa: 'mfa'
- note regex: input is JSON — use a JSON parser (jq, Logstash json filter, …) instead of a regex where possible
- note regex: a single linear template could not reproduce every input line — fields are captured with order-independent lookaheads instead
- note JSON input: the reused regex uses lookaheads (RE2 cannot compile them), so a native `- json:` stage is emitted instead — it extracts each field by key without a regex
syslog-ng
parser p_json {
regexp-parser(
prefix(".json.")
patterns("(?=.*?\"ts\":\"(?<ts>[^\"]*)\")(?=.*?\"level\":\"(?<level>[^\"]*)\")(?=.*?\"service\":\"(?<service>[^\"]*)\")(?=.*?\"msg\":\"(?<msg>[^\"]*)\")(?=.*?\"order_id\":\"(?<order_id>[^\"]*)\"|)(?=.*?\"user\":\"(?<user>[^\"]*)\")(?=.*?\"ip\":\"(?<ip>[^\"]*)\")(?=.*?\"name\":\"(?<gateway_name>[^\"]*)\"|)(?=.*?\"code\":\"(?<gateway_code>[^\"]*)\"|)(?=.*?\"mfa\":(?<mfa>[A-Za-z]+)|).*")
);
}; - note regex: input is JSON — use a JSON parser (jq, Logstash json filter, …) instead of a regex where possible
- note regex: a single linear template could not reproduce every input line — fields are captured with order-independent lookaheads instead
- note captured fields are stored as name-value pairs under the prefix ".json." (e.g. a group (?<srcip>…) becomes ".json.srcip")
- note json structure: syslog-ng has a dedicated json-parser() that handles nested objects natively — consider json-parser(prefix(".logforge.")) instead of the emitted regexp-parser
FAQ
- What is NDJSON / JSON Lines and why do logs use it?
- NDJSON (newline-delimited JSON), also called JSON Lines, is one complete JSON object per line with no enclosing array. Logs use it because each line is independently valid and parseable — a tail can stream them, a single corrupt line does not break the rest, and appending is trivial. A whole-file JSON array would require rewriting the closing bracket on every append.
- How are nested JSON log fields flattened for a SIEM?
- Nested objects are collapsed into dotted or underscored keys — gateway.name becomes gateway.name or gateway_name, and gateway.code becomes gateway.code / gateway_code. The exact delimiter depends on the shipper (Filebeat, Fluent Bit, Vector) and its config. Flattening lets flat-schema stores index and filter on what were nested paths.
- Why do my JSON log lines have different fields from each other?
- Because the schema is per-event: an error event may carry order_id, user, and an error object while an info event carries none of them. This is normal for structured logging. Downstream systems should treat the union of possible keys as optional, not require a fixed set — the presence or absence of a key is itself signal.
- Which key holds the timestamp in a JSON log?
- There is no universal key — it depends on the logging library. Common ones are ts, time, @timestamp (the Elastic convention), and t. Check what your framework emits; the value is usually ISO 8601 with milliseconds (2026-07-03T14:22:15.003Z), but some libraries write epoch seconds or milliseconds instead.
Try it on your own JSON application lines
Paste a few real lines, review the detected fields, and copy whichever format your stack needs. Free, no account, nothing uploaded.
Open this sample in LogForge →