SPARQL Grammar
The SPARQL grammar covers both SPARQL Query and SPARQL Update.
SPARQL Request String
A SPARQL Request String is a SPARQL Query String or SPARQL Update String and it is a Unicode character string (c.f. section 6.1 String concepts of [Character Model for the World Wide Web 1.0: Fundamentals]) in the language defined by the following grammar.
A SPARQL Query String start at the QueryUnit production.
A SPARQL Update String start at the UpdateUnit production.
For compatibility with future versions of Unicode, the characters in this string may
include Unicode codepoints that are unassigned as of the date of this publication
(see Identifier and Pattern Syntax [Identifier and Pattern Syntax 4.1.0] section 4 Pattern Syntax). For productions with excluded character classes (for
example [^<>'{}|^]), the characters are excluded from the range#x0 - #x10FFFF`.
Codepoint Escape Sequences
A SPARQL Query String is processed for codepoint escape sequences before parsing by the grammar defined in EBNF below. The codepoint escape sequences for a SPARQL query string are:
| Escape | Unicode code point |
|---|---|
| '\u'HEXHEXHEXHEX | A Unicode code point in the range U+0 to U+FFFF inclusive corresponding to the encoded hexadecimal value. |
| '\U'HEXHEXHEXHEXHEXHEXHEXHEX | A Unicode code point in the range U+0 to U+10FFFF inclusive corresponding to the encoded hexadecimal value. |
where HEX is a hexadecimal character
HEX ::= [0-9] | [A-F] | [a-f]
Examples:
<ab\u00E9xy> # Codepoint 00E9 is Latin small e with acute - é
\u03B1:a # Codepoint x03B1 is Greek small alpha - α
a\u003Ab # a:b -- codepoint x3A is colon
Codepoint escape sequences can appear anywhere in the query string. They are processed
before parsing based on the grammar rules and so may be replaced by codepoints with
significance in the grammar, such as ":" marking a prefixed name.
These escape sequences are not included in the grammar below. Only escape sequences
for characters that would be legal at that point in the grammar may be given. For
example, the variable "?x\u0020y" is not legal (\u0020 is a space and is not permitted in a variable name).
White Space
White space (production WS) is used to separate two terminals which would otherwise be (mis-)recognized as one terminal. Rule names below in capitals indicate where white space is significant; these form a possible choice of terminals for constructing a SPARQL parser. White space is significant in strings. Otherwise, white space is ignored between tokens.
For example:
?a<?b&&?c>?d
is the token sequence variable '?a', an IRI '<?b&&?c>', and variable '?d', not a expression involving the operator '&&' connecting two expression using '<' (less than) and '>' (greater than).
Comments
Comments in SPARQL queries take the form of '#', outside an IRI or string, and continue to the end of line (marked by characters
0x0D or 0x0A) or end of file if there is no end of line after the comment marker. Comments are
treated as white space.
IRI References
Text matched by the IRIREF production and PrefixedName (after prefix expansion) production,
after escape processing, must conform to the generic syntax of IRI references in section
2.2 of RFC 3987 "ABNF for IRI References and IRIs" [Internationalized Resource Identifiers (IRIs)]. For example, the IRIREF<abc#def> may occur in a SPARQL query string, but the IRIREF<abc##def> must not.
Base IRIs declared with the BASE keyword must be absolute IRIs. A prefix declared with the PREFIX keyword may not be re-declared in the same query. See section 4.1.1, Syntax of IRI Terms, for a description of BASE and PREFIX .
Blank Nodes and Blank Node Labels
Blank nodes can not be used in:
- DELETE WHERE
- DELETE DATA
- a DeleteClause
in a SPARQL Update request.
Blank node labels are scoped to the SPARQL Request String in which they occur. Different uses of the same blank node label in a request string refer to the same blank node. Fresh blank nodes are generated for each request; blank nodes can not be referenced by label across requests.
The same blank node label can not be used in:
- two basic graph patterns in a SPARQL Query
- two WHERE clauses within a single SPARQL Update request
- two INSERT DATA operations within a single SPARQL Update request
Note that the same blank node label can occur in different QuadPattern clauses in a SPARQL Update request.
Escape sequences in strings
In addition to the codepoint escape sequences, the following escape sequences any string production (e.g. STRING_LITERAL1, STRING_LITERAL2, STRING_LITERAL_LONG1, STRING_LITERAL_LONG2):
| Escape | Unicode code point |
|---|---|
| '\t' | U+0009 (tab) |
| '\n' | U+000A (line feed) |
| '\r' | U+000D (carriage return) |
| '\b' | U+0008 (backspace) |
| '\f' | U+000C (form feed) |
| '\"' | U+0022 (quotation mark, double quote mark) |
| "\'" | U+0027 (apostrophe-quote, single quote mark) |
| '\\' | U+005C (backslash) |
Examples:
"abc\n"
"xy\rz"
'xy\tz'
Grammar
The EBNF notation used in the grammar is defined in Extensible Markup Language (XML) 1.1 [Extensible Markup Language (XML) 1.1] section 6 Notation.
Notes:
- Keywords are matched in a case-insensitive manner with the exception of the keyword
'
a' which, in line with Turtle and N3, is used in place of the IRIrdf:type(in full, http://www.w3.org/1999/02/22-rdf-syntax-ns#type). - Escape sequences are case sensitive.
- When tokenizing the input and choosing grammar rules, the longest match is chosen.
- The SPARQL grammar is LL(1) when the rules with uppercased names are used as terminals.
- There are two entry points into the grammar:
QueryUnitfor SPARQL queries, andUpdateUnitfor SPARQL Update requests. - In signed numbers, no white space is allowed between the sign and the number. The AdditiveExpression grammar rule allows for this by covering the two cases of an expression followed by a signed number. These produce an addition or subtraction of the unsigned number as appropriate.
- The tokens
INSERT DATA, DELETE DATA,DELETE WHERE allow any amount of white space between the words. The single space version is used in the grammar for clarity. - The QuadData and QuadPattern rules both use rule Quads. The rule QuadData, used in
INSERT DATAandDELETE DATA, must not allow variables in the quad patterns. - Blank node syntax is not allowed in DELETE WHERE, the DeleteClause for
DELETE, nor inDELETE DATA. - Rules for limiting the use of blank node labels are given in section 19.6.
- The number of variables in the variable list of
VALUESblock must be the same as the number of each list of associated values in theDataBlock. - Variables introduced by
ASin aSELECTclause must not already be in-scope. - The variable assigned in a
BINDclause must not be already in-use within the immediately preceding TriplesBlock within a GroupGraphPattern. - Aggregate functions can be one of the built-in keywords for aggregates or a custom aggregate, which is syntactically a function call. Aggregate functions may only be used in SELECT, HAVING and ORDER BY clauses.
- Only custom aggregate functions use the DISTINCT keyword in a function call.
Productions for terminals:
[139] |
IRIREF |
::= | '<' ([^<>"{}|^`\]-[#x00-#x20])* '>' |
[140] |
PNAME_NS |
::= | PN_PREFIX? ':' |
[141] |
PNAME_LN |
::= | PNAME_NSPN_LOCAL |
[142] |
BLANK_NODE_LABEL |
::= | '_:' ( PN_CHARS_U | [0-9] ) ((PN_CHARS|'.')* PN_CHARS)? |
[143] |
VAR1 |
::= | '?' VARNAME |
[144] |
VAR2 |
::= | '$' VARNAME |
[145] |
LANGTAG |
::= | '@' [a-zA-Z]+ ('-' [a-zA-Z0-9]+)* |
[146] |
INTEGER |
::= | [0-9]+ |
[147] |
DECIMAL |
::= | [0-9]* '.' [0-9]+ |
[148] |
DOUBLE |
::= | [0-9]+ '.' [0-9]* EXPONENT | '.' ([0-9])+ EXPONENT | ([0-9])+ EXPONENT |
[149] |
INTEGER_POSITIVE |
::= | '+'INTEGER |
[150] |
DECIMAL_POSITIVE |
::= | '+'DECIMAL |
[151] |
DOUBLE_POSITIVE |
::= | '+'DOUBLE |
[152] |
INTEGER_NEGATIVE |
::= | '-'INTEGER |
[153] |
DECIMAL_NEGATIVE |
::= | '-'DECIMAL |
[154] |
DOUBLE_NEGATIVE |
::= | '-'DOUBLE |
[155] |
EXPONENT |
::= | [eE] [+-]? [0-9]+ |
[156] |
STRING_LITERAL1 |
::= | "'" ( ([^#x27#x5C#xA#xD]) | ECHAR )* "'" |
[157] |
STRING_LITERAL2 |
::= | '"' ( ([^#x22#x5C#xA#xD]) | ECHAR )* '"' |
[158] |
STRING_LITERAL_LONG1 |
::= | "'''" ( ( "'" | "''" )? ( [^'\] | ECHAR ) )* "'''" |
[159] |
STRING_LITERAL_LONG2 |
::= | '"""' ( ( '"' | '""' )? ( [^"\] | ECHAR ) )* '"""' |
[160] |
ECHAR |
::= | '\' [tbnrf\"'] |
[161] |
NIL |
::= | '(' WS* ')' |
[162] |
WS |
::= | #x20 | #x9 | #xD | #xA |
[163] |
ANON |
::= | '[' WS* ']' |
[164] |
PN_CHARS_BASE |
::= | [A-Z] | [a-z] | [#x00C0-#x00D6] | [#x00D8-#x00F6] | [#x00F8-#x02FF] | [#x0370-#x037D]
| [#x037F-#x1FFF] | [#x200C-#x200D] | [#x2070-#x218F] | [#x2C00-#x2FEF] | [#x3001-#xD7FF]
| [#xF900-#xFDCF] | [#xFDF0-#xFFFD] | [#x10000-#xEFFFF] |
[165] |
PN_CHARS_U |
::= | PN_CHARS_BASE | '_' |
[166] |
VARNAME |
::= | ( PN_CHARS_U | [0-9] ) ( PN_CHARS_U | [0-9] | #x00B7 | [#x0300-#x036F] | [#x203F-#x2040] )* |
[167] |
PN_CHARS |
::= | PN_CHARS_U | '-' | [0-9] | #x00B7 | [#x0300-#x036F] | [#x203F-#x2040] |
[168] |
PN_PREFIX |
::= | PN_CHARS_BASE ((PN_CHARS|'.')* PN_CHARS)? |
[169] |
PN_LOCAL |
::= | (PN_CHARS_U | ':' | [0-9] | PLX ) ((PN_CHARS | '.' | ':' | PLX)* (PN_CHARS | ':' | PLX) )? |
[170] |
PLX |
::= | PERCENT | PN_LOCAL_ESC |
[171] |
PERCENT |
::= | '%' HEXHEX |
[172] |
HEX |
::= | [0-9] | [A-F] | [a-f] |
[173] |
PN_LOCAL_ESC |
::= | '\' ( '_' | '~' | '.' | '-' | '!' | '$' | '&' | "'" | '(' | ')' | '*' | '+' | ','
| ';' | '=' | '/' | '?' | '#' | '@' | '%' ) |
Copyright © 2013 W3C® (MIT, ERCIM, Keio, Beihang). This software or document includes material copied from or derived from SPARQL 1.1 Query.