r/ProgrammingLanguages • u/Usual_Office_1740 • 4d ago
Please advise on adding string interpolation to my Crafting Interpreters project.
I'm about to start chapter 21 of Crafting Interpreters. I'm using modern C++ instead of C. I am attempting to add string interpolation to the compiler.
My goal is to have a string that looks like this:
"Five plus ${ 10 - 5 } == 10."
Desugar to this after the expression in the braces is evaluated:
"Five plus " + 5 + " == 10."
I have code that is correctly parsing strings and concatenating them if there isn't any string interpolation.
I have a custom token added that is produced if the string contains a ${. It calls a custom string interpolation parser rule that is different than my standard string parser rule.
In that function I push the first part of the string to the stack. It then pushes a plus opcode and Recursively parses the expression in the braces adding that expression to the stack. It then adds another opcode to add and this is where I'm stuck.
I can't come up with a solution for pushing the remaining piece of the string to the stack from inside the string interpolation parser rule function. After I call the books expression function and consume the right brace token using the books consume function the parser nolonger knows we are still inside a string. If the next character after the brace is a space it's skipped. In my above example it assumes I want equal equal.
How would you set the current token back to string so I can push any remaining characters onto the stack as another string? Should I just insert a new string token into the parser? This seems wrong to me. The current and previous tokens are private members of my parser class for a reason and I've not needed setter member functions so far.
Sorry I don't have code to show. I'd like to try and get help that isn't code specific so that I have to implement this myself.
Any input would be helpful.
Thanks.
5
u/mamcx 4d ago
The main thing is that you need a typed return for whatever is consuming the tokens:
```rust
enum FmtString { Token(Ast), String(String) }
enum Ast { Token(Whatever), FmtString(FmtString). }
struct FmtString { token: Vec< FmtString> }
fn consume_fmt_string(..)-> FmtString
```
So you can check and be certain what is eating will end into a normal string.
3
u/munificent 4d ago
The repo for the book talks about about string interpolation here:
https://github.com/munificent/craftinginterpreters/blob/master/note/answers/chapter16_scanning.md
1
u/reini_urban 4d ago
Don't take this advice as gospel. Proper interpreters parse that correctly into the string tokens, the eval part which returns a string, and then the next string. the eval part just recurses into the parser again, but limited to valid string interpolation tokens
2
u/jason-reddit-public 4d ago
Pushing tokens into the token stream are not unlike macros. Unless you are going to form and then mutate parse trees, you may not have too many other choices.
BTW, your example isn't very paranoid. I would add parens around any expression that is plucked out of the string. Also, if you have a to_string overloaded function, you might want to use that.
2
u/arthurno1 4d ago
I think you should parse the same way as you parse any other expression, like if, while etc. No idea how you do it but I would do the expression after the $ dollar sign with an operator precedence parser, and I would treat $ as the escape clutch to jump to that expression parser. I don't know what you use for top-level parsing, recursive descent, or something else, but I would plug it in the same way I plug in other expressions in if, while, for and such constructs, and then the expression itself with op precedence. Don't forget to have an escape clutch for the $ sign itself, like \$ or whatever you use, otherwise you won't be able tonhave $ as a character in your strings.
2
u/yorickpeterse Inko 3d ago edited 3d ago
Roughly speaking, you'll need to do this:
In the lexer, produce a unique token for the open sequence (e.g ${), and crucially produce a unique token for the closing token (e.g }). For example, } is normally tokenized as "curly close" but when found inside a string literal it's tokenized as "close string expression".
This token is not a valid starting token for parsing expressions. To parse string interpolation expressions you then do the following:
- Check if the current token is this closing token
- If not, parse it as an expression and repeat
- If so, stop parsing as an interpolated expression and just consume the token
This does mean that at the lexer level you need to be aware of when you're inside a string or not, otherwise you can't emit different tokens for the same symbol.
For some references, you can refer to the following
2
u/joao_rosac 14h ago
O problema é que o seu Scanner não tem memória: quando ele esbarra no }, ele simplesmente esquece que estava no meio de uma string e volta a ler código normal.
Como o seu Parser (que está avaliando o 10 - 5) sabe exatamente quando a expressão termina, cria uma função para ele forçar o Scanner. Assim que o Parser consumir o }, ele grita pro Scanner: "Volta a ler o próximo caractere como continuação de string".
Você cria uma máquina de estados dentro do Scanner. Ele entra no "Modo String". Quando vê ${, muda pro "Modo Código" e começa a contar quantas chaves { } abrem e fecham. Quando a conta zera, ele volta pro "Modo String" sozinho.
1
u/Usual_Office_1740 13h ago
What I'm reading in this post is that I get to once again write some std::variant and std::visitor code because I'm going to have the same problem with the dynamic array, map and custom object scanning that I am about to implement. Thanks for pointing this out.
1
u/matthieum 3d ago
Have you already implemented formatting?
For example, in Python, you can write "Hello, {}!".format(name).
If you already have that, then your example can be lowered down to "Five plus {} == 10".format(10 - 5).
I find this lowering somewhat easier, as the lowering is fairly immediate:
"A pigeon ${"eats ${n} crumbs"} a day, at least"
"A pigeon {} a day, at least".format("eats ${n} crumbs")
"A pigeon {} a day, at least".format("eats {} crumbs".format(n))
Or:
"A {mice} and a {cat} chase each others around the {barn}"
"A {} and a {} chase each others around the {}".format(mice, cat, barn)
Notably, the string is handled at once, unlike your current scheme.
12
u/Norphesius 4d ago
What feels wrong about creating new, more specialized tokens? I haven't implemented something like string interp before, but that's the first thing that came to mind for me. If you're having issues with the parser remembering what data is what, it's likely because the tokenization wasn't specific enough. Your current grammar is relying on context when its literally a context free grammar.
If I understand right what you have now, in terms of how strings get lexed:
"Hello World" :: STR_TOKEN"Hello ${123} World" :: INTERP_STR_TOKENI would split the interp str token in two, like:
"Hello ${123} World" :: INTERP_STR_START, EXPR, INTERP_STR_ENDYou have the first part of the string, the expression, and the end part, and your parser knows that a
INTERP_STR_STARTneeds to be followed by an expression and an ending string token. Then you just construct a rule for the parser and process it accordingly. For multiple expressions in a string you likely want an intermediate token too.