What an Apache access (combined) line looks like
The Combined sample below is fed verbatim into the engine to produce every parser on this page.
192.0.2.10 - jdoe [03/Jul/2026:14:22:15 +0300] "GET /wp-admin/ HTTP/1.1" 302 512 "https://example.com/" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/126.0"
203.0.113.99 - - [03/Jul/2026:14:22:40 +0300] "GET /.env HTTP/1.1" 404 153 "-" "curl/8.6.0" Detected fields
The engine classified this sample as freeform and consolidated 11 fields across 2 lines. Fields marked literal were identical on every sample line, so they are baked into the pattern as anchors rather than captured.
- ip1 : ipv4
- _lit1 : literal · literal
- literal : literal
- timestamp : timestamp
- method : http_method · literal
- quoted_string : quoted_string
- quoted_string2 : quoted_string · literal
- status : http_status
- number : number
- url : url
- user_agent : user_agent
Regex (named capture groups)
# sample: 192.0.2.10 - jdoe [03/Jul/2026:14:22:15 +0300] "GET /wp-admin/ HTTP/1.1" 302 512 "https://example.com/" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/126.0"
# groups: ip1=192.0.2.10, literal=jdoe, timestamp=03/Jul/2026:14:22:15 +0300, quoted_string=/wp-admin/, status=302, number=512, url=https://example.com/, user_agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/126.0
^(?<ip1>\d{1,3}(?:\.\d{1,3}){3}) - (?<literal>(?:[A-Za-z]+|-)) \[(?<timestamp>\d+/[A-Za-z]+/\d+:\d+:\d+:\d+ \+\d+)\] "GET (?<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\.[A-Za-z]+)) HTTP/1\.1" (?<status>\d{3}) (?<number>-?\d+(?:\.\d+)?) "(?<url>[^"]*)" "(?<user_agent>[^"]*)"$ - note the message tail diverges across lines and could not be split into stable fields — the trailing per-token literal group(s) are positional word-slices of that unstructured message, not stable fields
Grok pattern (Logstash / Elastic)
# custom patterns
APACHE_NOTDQUOTE [^"]*
%{IPV4:ip1} - %{NOTSPACE:literal} \[%{HTTPDATE:timestamp}\] "GET %{NOTSPACE:quoted_string} HTTP/1\.1" %{INT:status} %{NUMBER:number} "%{APACHE_NOTDQUOTE:url}" "%{APACHE_NOTDQUOTE:user_agent} - note constant field "method" embedded as literal anchor "GET" (varying=false)
- note constant field "quoted_string2" embedded as literal anchor "HTTP/1.1" (varying=false)
- note field "url" (url): samples do not all match %{URI}; using %{APACHE_NOTDQUOTE} instead
- note custom patterns emitted — save the '# custom patterns' block to a file in your patterns_dir
Wazuh decoder (OS_Regex XML)
<!--
Generated by LogForge - Wazuh decoder (OS_Regex dialect, not PCRE)
sample: 192.0.2.10 - jdoe [03/Jul/2026:14:22:15 +0300] "GET /wp-admin/ HTTP/1.1" 302 512 "https://example.com/" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/5
test with: /var/ossec/bin/wazuh-logtest
-->
<decoder name="apache-freeform">
<prematch>^\d+.\d+.\d+.\d+ </prematch>
</decoder>
<decoder name="apache-freeform">
<parent>apache-freeform</parent>
<regex>^(\d+.\d+.\d+.\d+) - (\w+) [(\d+/\w+/\d+:\d+:\d+:\d+ \p\d+)] "GET (\S+) HTTP/1.1" (\d+) (\d+) "(\.+)" "(\.+)"</regex>
<order>srcip, literal, timestamp, quoted_string, status, number, url, user_agent</order>
</decoder>
<!-- ============================================================
ALERT RULE (starter) — put this in a RULES file, e.g.
/var/ossec/etc/rules/local_rules.xml. Decoders and rules live
in SEPARATE files. The rule matches the decoder above through
<decoded_as>; set <level> and add <field>/<match> conditions so
it alerts only on the events you care about. Rule ids 100000+
are the user range — change them if they collide with yours.
============================================================ -->
<group name="apache,">
<rule id="100000" level="3">
<decoded_as>apache-freeform</decoded_as>
<description>apache: srcip=$(srcip) url=$(url) status=$(status)</description>
</rule>
<!-- Example — a higher-level alert gated on one field (uncomment and edit):
<rule id="100001" level="10">
<if_sid>100000</if_sid>
<field name="status">^5\d\d$</field>
<description>apache: a server-error response from $(status)</description>
</rule>
-->
</group>
- note no stable literal prefix found — <prematch> anchors on the leading field pattern; tighten it for your environment
- note field "ip1" mapped to Wazuh conventional field "srcip"
- note field "url": free-text capture (\.+) bounded by a quote anchor — OS_Regex greediness may over-consume if the anchor repeats
- note field "user_agent": free-text capture (\.+) bounded by end of line — OS_Regex greediness may over-consume if the anchor repeats
- note constant field "method" embedded as literal anchor "GET"
- note constant field "quoted_string2" embedded as literal anchor "HTTP/1.1"
- note added a starter alert <rule> (level 3, matched to the decoder via <decoded_as>) — put it in a RULES file (not the decoders file), set the level, and add <field>/<match> conditions; the commented example child rule shows the pattern
- note decoder order and prematch specificity may need site-specific tuning (other decoders in your ruleset can shadow these) — validate with /var/ossec/bin/wazuh-logtest
Wazuh's OS_Regex is not PCRE — a bare . is a literal dot and \. matches any character.
Test Wazuh OS_Regex patterns →
rsyslog template / liblognorm rulebase
version=2
# apache — liblognorm v2 rulebase (generated by LogForge)
# Usage with rsyslog (mmnormalize runs liblognorm):
# module(load="mmnormalize")
# action(type="mmnormalize" rulebase="/etc/rsyslog.d/apache.rb" useRawMsg="on")
# Literal "%" is escaped as "%%"; raw tabs are written as \x09.
rule=apache:%ip1:ipv4% - %literal:word% [%timestamp:char-to{"extradata":"]"}%] "GET %quoted_string:word% HTTP/1.1" %status:number% %number:number% "%url:char-to{"extradata":"\""}%" "%user_agent:char-to{"extradata":"\""}%"
- note trailing literal "\"" reconstructed from line 1
- note field "timestamp": samples do not uniformly match engine type "timestamp"; using a generic parser
- note chosen parser types: ip1=ipv4, literal=word, timestamp=char-to(]), quoted_string=word, status=number, number=number, url=char-to("), user_agent=char-to(")
Splunk
# props.conf (search-time extraction)
[<REPLACE_WITH_SOURCETYPE>]
EXTRACT-logforge = (?<ip1>\d{1,3}(?:\.\d{1,3}){3}) - (?<literal>(?:[A-Za-z]+|-)) \[(?<timestamp>\d+/[A-Za-z]+/\d+:\d+:\d+:\d+ \+\d+)\] "GET (?<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\.[A-Za-z]+)) HTTP/1\.1" (?<status>\d{3}) (?<number>-?\d+(?:\.\d+)?) "(?<url>[^"]*)" "(?<user_agent>[^"]*)"
# Quick search-time test in SPL:
# | rex field=_raw "(?<ip1>\\d{1,3}(?:\\.\\d{1,3}){3}) - (?<literal>(?:[A-Za-z]+|-)) \\[(?<timestamp>\\d+/[A-Za-z]+/\\d+:\\d+:\\d+:\\d+ \\+\\d+)\\] \"GET (?<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\\.[A-Za-z]+)) HTTP/1\\.1\" (?<status>\\d{3}) (?<number>-?\\d+(?:\\.\\d+)?) \"(?<url>[^\"]*)\" \"(?<user_agent>[^\"]*)\"" - note regex: the message tail diverges across lines and could not be split into stable fields — the trailing per-token literal group(s) are positional word-slices of that unstructured message, not stable fields
- note EXTRACT-<class> names must be unique within a sourcetype stanza — rename EXTRACT-logforge if you already use that class for this sourcetype
- note a timestamp field was detected: this EXTRACT only makes it a searchable field. To set the event _time at index time, add TIME_PREFIX and TIME_FORMAT to this props.conf stanza (TIME_FORMAT uses Splunk strptime, e.g. %Y-%m-%dT%H:%M:%S) — this generator does not guess the strptime format.
ES ingest
PUT _ingest/pipeline/apache
{
"description": "LogForge-generated ingest pipeline for apache",
"processors": [
{
"grok": {
"field": "message",
"patterns": [
"%{IPV4:ip1} - %{NOTSPACE:literal} \\[%{HTTPDATE:timestamp}\\] \"GET %{NOTSPACE:quoted_string} HTTP/1\\.1\" %{INT:status} %{NUMBER:number} \"%{APACHE_NOTDQUOTE:url}\" \"%{APACHE_NOTDQUOTE:user_agent}"
],
"pattern_definitions": {
"APACHE_NOTDQUOTE": "[^\"]*"
}
}
}
]
} - note grok: constant field "method" embedded as literal anchor "GET" (varying=false)
- note grok: constant field "quoted_string2" embedded as literal anchor "HTTP/1.1" (varying=false)
- note grok: field "url" (url): samples do not all match %{URI}; using %{APACHE_NOTDQUOTE} instead
- note grok: custom patterns emitted — save the '# custom patterns' block to a file in your patterns_dir
- note test in Kibana Dev Tools with: POST _ingest/pipeline/apache/_simulate (supply a docs[] array whose _source.message holds a sample line)
Graylog
# Grok patterns to add under System > Grok Patterns:
# (Graylog needs these custom patterns installed globally BEFORE the rule/extractor below will work.)
# APACHE_NOTDQUOTE [^"]*
# --- Graylog processing pipeline rule (primary) ---
# Paste under System > Pipelines > Manage rules, then attach the rule to a pipeline stage.
rule "apache-parse"
when
has_field("message")
then
let gp = grok(pattern: "%{IPV4:ip1} - %{NOTSPACE:literal} \\[%{HTTPDATE:timestamp}\\] \"GET %{NOTSPACE:quoted_string} HTTP/1\\.1\" %{INT:status} %{NUMBER:number} \"%{APACHE_NOTDQUOTE:url}\" \"%{APACHE_NOTDQUOTE:user_agent}", value: to_string($message.message), only_named_captures: true);
set_fields(gp);
end
# --- Graylog import-ready extractor JSON (secondary) ---
# Save as a .json file and import under System > Inputs > (input) > Manage extractors > Actions > Import extractors.
{
"extractors": [
{
"title": "apache",
"extractor_type": "grok",
"converters": [],
"order": 0,
"cursor_strategy": "copy",
"source_field": "message",
"target_field": "",
"extractor_config": {
"grok_pattern": "%{IPV4:ip1} - %{NOTSPACE:literal} \\[%{HTTPDATE:timestamp}\\] \"GET %{NOTSPACE:quoted_string} HTTP/1\\.1\" %{INT:status} %{NUMBER:number} \"%{APACHE_NOTDQUOTE:url}\" \"%{APACHE_NOTDQUOTE:user_agent}",
"named_captures_only": true
},
"condition_type": "none",
"condition_value": ""
}
],
"version": "5.0.0"
} - note grok: constant field "method" embedded as literal anchor "GET" (varying=false)
- note grok: constant field "quoted_string2" embedded as literal anchor "HTTP/1.1" (varying=false)
- note grok: field "url" (url): samples do not all match %{URI}; using %{APACHE_NOTDQUOTE} instead
- note grok: custom patterns emitted — save the '# custom patterns' block to a file in your patterns_dir
- note 1 custom grok pattern(s) (APACHE_NOTDQUOTE) must be installed globally first under System > Grok Patterns — see the block at the top of the output
- note primary artifact is the processing-pipeline rule; the extractor JSON is an equivalent import-ready alternative for the classic extractor UI
Datadog
logforge_rule %{ipv4:ip1} - %{notSpace:literal} \[%{date("dd/MMM/yyyy:HH:mm:ss Z"):timestamp}\] "GET %{quotedString:quoted_string} HTTP/1\.1" %{integer:status} %{number:number} "%{notSpace:url}" "%{data:user_agent} - note emitted rule name is "logforge_rule"; rename it to match your "apache" convention if desired
- note constant field "method" embedded as literal anchor "GET" (varying=false)
- note constant field "quoted_string2" embedded as literal anchor "HTTP/1.1" (varying=false)
- note paste this line into a Grok Parser processor in a Datadog Log Pipeline; matchers are anchored left-to-right and rule whitespace matches log whitespace. Complex or multi-shape logs may need Helper Rules.
Fluent Bit
[PARSER]
Name apache
Format regex
Regex ^(?<ip1>\d{1,3}(?:\.\d{1,3}){3}) - (?<literal>(?:[A-Za-z]+|-)) \[(?<timestamp>\d+/[A-Za-z]+/\d+:\d+:\d+:\d+ \+\d+)\] "GET (?<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\.[A-Za-z]+)) HTTP/1\.1" (?<status>\d{3}) (?<number>-?\d+(?:\.\d+)?) "(?<url>[^"]*)" "(?<user_agent>[^"]*)"$
Time_Key timestamp
Time_Format %d/%b/%Y:%H:%M:%S %z
# Fluentd <parse> block:
# <parse>
# @type regexp
# expression /^(?<ip1>\d{1,3}(?:\.\d{1,3}){3}) - (?<literal>(?:[A-Za-z]+|-)) \[(?<timestamp>\d+\/[A-Za-z]+\/\d+:\d+:\d+:\d+ \+\d+)\] "GET (?<quoted_string>(?:\/[A-Za-z]+-[A-Za-z]+\/|\/\.[A-Za-z]+)) HTTP\/1\.1" (?<status>\d{3}) (?<number>-?\d+(?:\.\d+)?) "(?<url>[^"]*)" "(?<user_agent>[^"]*)"$/
# time_key timestamp
# time_format %d/%b/%Y:%H:%M:%S %z
# </parse>
- note regex: the message tail diverges across lines and could not be split into stable fields — the trailing per-token literal group(s) are positional word-slices of that unstructured message, not stable fields
- note Time_Key set to "timestamp"; Time_Format "%d/%b/%Y:%H:%M:%S %z" is a best-effort strptime derived from the sample shape — verify it against your data (Fluent Bit uses %L for fractional seconds and %z for numeric offsets)
Vector
[transforms.apache_parse]
type = "remap"
inputs = ["REPLACE_WITH_SOURCE"]
source = '''
. |= parse_regex!(.message, r'(?P<ip1>\d{1,3}(?:\.\d{1,3}){3}) - (?P<literal>(?:[A-Za-z]+|-)) \[(?P<timestamp>\d+/[A-Za-z]+/\d+:\d+:\d+:\d+ \+\d+)\] "GET (?P<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\.[A-Za-z]+)) HTTP/1\.1" (?P<status>\d{3}) (?P<number>-?\d+(?:\.\d+)?) "(?P<url>[^"]*)" "(?P<user_agent>[^"]*)"')
''' - note [regex] the message tail diverges across lines and could not be split into stable fields — the trailing per-token literal group(s) are positional word-slices of that unstructured message, not stable fields
Loki
# promtail pipeline for "apache" (generated by LogForge)
# Add these stages under a scrape_config in your promtail config:
# scrape_configs:
# - job_name: apache
# pipeline_stages:
# (the stages below are indented to sit under pipeline_stages)
pipeline_stages:
- regex:
expression: '^(?P<ip1>\d{1,3}(?:\.\d{1,3}){3}) - (?P<literal>(?:[A-Za-z]+|-)) \[(?P<timestamp>\d+/[A-Za-z]+/\d+:\d+:\d+:\d+ \+\d+)\] "GET (?P<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\.[A-Za-z]+)) HTTP/1\.1" (?P<status>\d{3}) (?P<number>-?\d+(?:\.\d+)?) "(?P<url>[^"]*)" "(?P<user_agent>[^"]*)"$'
- labels:
status:
- note regex: the message tail diverges across lines and could not be split into stable fields — the trailing per-token literal group(s) are positional word-slices of that unstructured message, not stable fields
- note promoted low-cardinality field(s) to Loki labels: status
- note left in the extracted map (NOT promoted to labels — high cardinality would explode Loki streams): ip1, literal, timestamp, quoted_string, number, url, user_agent
syslog-ng
parser p_apache {
regexp-parser(
prefix(".apache.")
patterns("(?<ip1>\\d{1,3}(?:\\.\\d{1,3}){3}) - (?<literal>(?:[A-Za-z]+|-)) \\[(?<timestamp>\\d+/[A-Za-z]+/\\d+:\\d+:\\d+:\\d+ \\+\\d+)\\] \"GET (?<quoted_string>(?:/[A-Za-z]+-[A-Za-z]+/|/\\.[A-Za-z]+)) HTTP/1\\.1\" (?<status>\\d{3}) (?<number>-?\\d+(?:\\.\\d+)?) \"(?<url>[^\"]*)\" \"(?<user_agent>[^\"]*)\"")
);
}; - note regex: the message tail diverges across lines and could not be split into stable fields — the trailing per-token literal group(s) are positional word-slices of that unstructured message, not stable fields
- note captured fields are stored as name-value pairs under the prefix ".apache." (e.g. a group (?<srcip>…) becomes ".apache.srcip")
FAQ
- What is the difference between Apache combined and common log format?
- The 'common' format (CLF) ends after the response size: host, ident, user, time, request, status, bytes. 'combined' appends two quoted header fields, Referer and User-Agent. Both are just nicknames defined by LogFormat directives, so a given server logs whatever its active LogFormat says — inspect the config, not the file name.
- Why does my Apache line have an extra field my parser does not expect?
- Almost certainly a customized LogFormat. Common additions are a leading %v (virtual host), a %D/%T response-time column, %p (port), or %{X-Forwarded-For}i to capture the real client behind a proxy. Grab a representative sample from the actual server and regenerate the parser against it rather than assuming stock combined.
- How do I recover the real client IP when Apache is behind a load balancer?
- The %h field will be the proxy or load-balancer address. To get the originating client you need the X-Forwarded-For header, which is only present if the LogFormat includes %{X-Forwarded-For}i. If it does, capture that field and take the left-most address in the comma-separated list as the client (trusting it only as far as you trust your proxy chain).
Try it on your own Apache access (combined) lines
Paste a few real lines, review the detected fields, and copy whichever format your stack needs. Free, no account, nothing uploaded.
Open this sample in LogForge →