Module 3 · Strings and Template Literals
When to Use Which: Methods or a Loop, a Template or +, and When It Is a Regex
In this lesson
- Choose
splitandjoinwhen the text has a clean separator, and a character loop with aninQuotesflag when it does not. - Pick
sliceoversubstringandsubstr,includesorindexOf, a template literal,+orjoin, and<orlocaleCompare, and give the reason for each. - Use the chart and the flowchart to take text apart, keep an emoji whole with
[...s], and recognise when a job has become a regular expression.
Kenji imports the chess club's member list from a spreadsheet. He wrote the importer in ten minutes and made it fast: one split(",") per line. Then the club's treasurer asks why Amara lives in a town called Lake Road". Her address, 12, Lake Road, held a comma, and every column after it moved one place right. Fast and wrong is still wrong. This lesson finds the point where split stops working, then makes six more judgement calls, each with a program.
Where split stops working
A CSV file puts a field in double quotes when the field holds a comma. That is the rule from the CSV standard, RFC 4180. split(",") knows nothing about quotes, so it cuts inside them.
const record = "Kenji,\"12, Lake Road\",Dhaka";
const fields = record.split(",");
console.log(fields.length);
console.log(fields);
4
[ 'Kenji', '"12', ' Lake Road"', 'Dhaka' ]
Three fields went in and four came out, with the quote marks stuck to the pieces. No error appears, so the bug travels into the spreadsheet. split is the right tool only when the separator can never appear inside a field. Here it can, so the text needs a reader that remembers whether it is inside quotes.
A loop that knows it is inside quotes
A flag is a boolean variable that remembers one fact while a loop runs. Here the fact is "am I inside quotes?". A quote mark flips the flag. A comma ends a field only when the flag is off. Every other character joins the current field.
The quote-aware loop
let inQuotes = false;
for (const ch of record) {
if (ch === "\"") inQuotes = !inQuotes;
else if (ch === "," && !inQuotes) end the field
else add ch to the field
}
inQuotesstartsfalse, because a record starts outside quotes."\""is a string holding one double quote; the backslash keeps it from closing the string.!inQuotesis true only outside quotes, so a comma inside them is just text.
The program reads one record and prints its fields one per line, with split's count beside the loop's.
const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];
const record = next();
const fields = [];
let field = "";
let inQuotes = false;
for (const ch of record) {
if (ch === "\"") {
inQuotes = !inQuotes;
} else if (ch === "," && !inQuotes) {
fields.push(field);
field = "";
} else {
field += ch;
}
}
fields.push(field);
out.push(`split(","): ${record.split(",").length} pieces`);
out.push(`the loop: ${fields.length} fields`);
for (let i = 0; i < fields.length; i++) {
out.push(`${i + 1}: ${fields[i]}`);
}
console.log(out.join("\n"));
split(","): 4 pieces
the loop: 3 fields
1: Kenji
2: 12, Lake Road
3: Dhaka
That output is for the input Kenji,"12, Lake Road",Dhaka. The comma inside the quotes stayed in the address, and the quote marks themselves were dropped. The last push after the loop saves the final field, which no comma ends.
One honest gap remains. RFC 4180 writes a quote inside a quoted field as two quotes, "". This loop flips the flag twice and loses both, so "She said ""hi""" comes out as She said hi. Exercise 3 asks you to keep it.
Here is the whole decision as one picture. Start at the top, answer each question, and stop at the first yes.
slice, substring and substr: the chart
Lesson 02 met three ways to cut a piece out of a string. They agree on the easy calls and disagree on the rest. The chart runs all three on the same six calls, so the difference is visible.
Each line prints the three answers for one row, in brackets, so an empty string still shows.
const s = "template";
console.log(`(1, 4): [${s.slice(1, 4)}] [${s.substring(1, 4)}] [${s.substr(1, 4)}]`);
console.log(`(4, 1): [${s.slice(4, 1)}] [${s.substring(4, 1)}] [${s.substr(4, 1)}]`);
console.log(`(-3): [${s.slice(-3)}] [${s.substring(-3)}] [${s.substr(-3)}]`);
console.log(`(2, -1): [${s.slice(2, -1)}] [${s.substring(2, -1)}] [${s.substr(2, -1)}]`);
console.log(`(-4, 2): [${s.slice(-4, 2)}] [${s.substring(-4, 2)}] [${s.substr(-4, 2)}]`);
console.log(`(3): [${s.slice(3)}] [${s.substring(3)}] [${s.substr(3)}]`);
(1, 4): [emp] [emp] [empl]
(4, 1): [] [emp] [l]
(-3): [ate] [template] [ate]
(2, -1): [mplat] [te] []
(-4, 2): [] [te] [la]
(3): [plate] [plate] [plate]
Row (4, 1) is Bob's off-by-one waiting to happen. slice says "nothing there" and the empty result shows up in a test. substring quietly swaps the two numbers, and the wrong piece looks fine. So write slice: negatives count from the end, the end is excluded, and nothing is swapped behind your back.
includes for "is it there", indexOf for "where"
Both search a string. includes answers a yes-or-no question with a boolean. indexOf answers "where?" with a number, and -1 for "nowhere". Pick by the question you are asking.
Is there an @? Where is it? And what is the name in front of it?
const email = "kenji@mail.com";
console.log(email.includes("@"));
console.log(email.indexOf("@"));
console.log(email.slice(0, email.indexOf("@")));
console.log(email.includes("#"), email.indexOf("#"));
true
5
kenji
false -1
For "is it there", includes reads like the question and cannot fall into lesson 02's trap, where if (s.indexOf(x)) treats a match at index 0 as false. When you need the position, as for the name, indexOf is the tool, and you compare its answer with -1.
A template for a line with values, + for two pieces, join for many
Three ways build text, and each has its job. A template literal shows the finished line with holes in it, so you read it as it will print. + is fine for two pieces. join glues a whole array with one separator.
The same sentence with + and as a template, then a short label, then a CSV row from an array.
const name = "Amara";
const books = 3;
const total = 1940;
console.log("Dear " + name + ", your " + books + " books cost Tk " + total + ".");
console.log(`Dear ${name}, your ${books} books cost Tk ${total}.`);
console.log("Tk " + total);
const row = ["Amara", "16", "Dhaka", "amara@mail.com"];
console.log(row.join(","));
Dear Amara, your 3 books cost Tk 1940.
Dear Amara, your 3 books cost Tk 1940.
Tk 1940
Amara,16,Dhaka,amara@mail.com
The first two lines print the same text. In the + line, count the quote marks and the spaces you had to place by hand; one missing space and the receipt says 3books. For two pieces, "Tk " + total is short and clear. For a list, join puts exactly one separator between items and none at the ends.
Many teams make the template a rule. ESLint's prefer-template rule flags + between text and a value, and it can rewrite the line for you. That is a style choice that a team agrees on, not a speed rule: both lines produce the same string.
< for codes, localeCompare or Intl.Collator for people
Lesson 02 showed that < compares strings code by code, so every capital comes before every small letter. That is right for codes and IDs. It is wrong for a list of names a person will read. localeCompare compares the way a language orders words. An Intl.Collator is an object that does the same comparison with its settings fixed once, which suits a long sort.
Members typed their own names, some with a capital and one with an accent. The default sort() compares as text in code order. The arrow function inside sort is the short form Module 5 explains.
const names = ["zara", "Amara", "Émile", "bob", "Kenji"];
names.sort();
console.log(`< order: ${names.join(", ")}`);
names.sort((x, y) => x.localeCompare(y));
console.log(`localeCompare: ${names.join(", ")}`);
const collator = new Intl.Collator("en");
names.sort(collator.compare);
console.log(`Intl.Collator: ${names.join(", ")}`);
const photos = ["photo10.jpg", "photo9.jpg", "photo2.jpg"];
photos.sort(new Intl.Collator("en", { numeric: true }).compare);
console.log(`numeric: true: ${photos.join(", ")}`);
< order: Amara, Kenji, bob, zara, Émile
localeCompare: Amara, bob, Émile, Kenji, zara
Intl.Collator: Amara, bob, Émile, Kenji, zara
numeric: true: photo2.jpg, photo9.jpg, photo10.jpg
In code order, bob waits behind Kenji, and Émile lands last, because É is code 201. localeCompare and the collator give the order a reader expects. The option numeric: true compares runs of digits as numbers, so photo10 comes after photo9. For Bangla names, new Intl.Collator("bn") uses the Bangla rules.
[...s] over split("") when an emoji may appear
split("") cuts between code units, and lesson 01 showed that an emoji takes two. The spread [...s] and Array.from(s) walk the string by code point instead, so an emoji stays whole. Module 7 shows spread on arrays.
JSON.stringify prints each piece in quotes and writes a lone half of an emoji as an escape such as \ud83d. Printed directly, a lone half shows as a replacement character instead.
const name = "Zara 😀";
console.log(name.length);
console.log(JSON.stringify(name.split("")));
console.log(JSON.stringify([...name]));
console.log(JSON.stringify(Array.from(name)));
console.log(JSON.stringify(name.split("").reverse().join("")));
console.log([...name].reverse().join(""));
7
["Z","a","r","a"," ","\ud83d","\ude00"]
["Z","a","r","a"," ","😀"]
["Z","a","r","a"," ","😀"]
"\ude00\ud83d araZ"
😀 araZ
split("") gave seven pieces, the last two half an emoji each, and reversing put the halves in the wrong order. [...name] gave six pieces and a clean reverse. For text that is ASCII by its rules, such as the judge's words, split("") is fine. For anything a person typed, use [...s]. Bangla vowel signs still split from their letters, and that is lesson 05's Intl.Segmenter.
When the job has become a regular expression
Every tool above looks for fixed text. Some jobs look for a pattern: any run of spaces, any digit, a line made only of digits. That is the job of a regular expression, a small language for describing text, which Module 16 teaches. Here are two one-liners, named and not taught.
console.log("too many spaces".replace(/\s+/g, " "));
console.log(/^\d+$/.test("2026"), /^\d+$/.test("20x6"));
too many spaces
true false
The first squeezes every run of spaces into one. The second asks whether a text is digits from start to end. Each would take a loop of several lines. The sign is the word "any": when you hear yourself say "any run of" or "any digit", the job has become a pattern.
Where this is used
- ESLint. The rule
prefer-template("Require template literals instead of string concatenation") flags"Hello, " + name + "!". It is not in ESLint's recommended set, so a team turns it on by choice, and--fixrewrites the line. - The Airbnb JavaScript Style Guide. Its rule 6.3 says to use template strings instead of concatenation when building up strings, and names
prefer-templateandtemplate-curly-spacingas the ESLint rules. - RFC 4180. The 2005 CSV standard. Its rule 6 puts a field holding a comma in double quotes, which Example 1 handles. Its rule 7 writes a quote inside a field as two quotes, which Exercise 3 handles.
- ICU. Node and Chrome build
Intl.Collatoron ICU, the International Components for Unicode library, which follows the Unicode Collation Algorithm. Android hands the same library to apps asandroid.icu.text.Collator.
Common mistakes
1. Reading substr's second number as an end.
console.log("hello".substr(1, 3), "hello".slice(1, 3));
ell el
substr(1, 3) means three characters from index 1, while slice(1, 3) stops before index 3. No error appears, just one character too many. Write slice. You will meet substr in old code and assume it matches its neighbours.
2. Using includes where a position is needed.
const mail = "kenji@mail.com";
const at = mail.includes("@");
console.log(at);
console.log(mail.slice(at));
true
enji@mail.com
slice wanted a number, got true, and converted it to 1, as Module 2's coercion rules say. The program cut one letter and kept going. Use indexOf when you need where. You will reach for includes because it is the newer, friendlier name.
3. Adding numbers inside a + chain.
console.log("Total: " + 2 + 3);
console.log(`Total: ${2 + 3}`);
Total: 23
Total: 5
+ works left to right, so the text meets 2 first and every + after it joins text. Inside ${}, the sum is worked out on its own. You will write the first line because it worked a moment ago with one number.
Kenji's club keeps every member's email on one line, separated by commas. The secretary types a name to look for, and wants to know whether it is there and where.
Input. Line 1 is the list. Line 2 is n. Then n lines, each a text to look for.
Output. One line per text: found at and the index of its first match, or not found.
Constraints. 1 <= n <= 1000. The list has 1 to 2000 characters; each text has 1 to 50.
Sample. The list zara@mail.com,bob@mail.com,kenji@mail.com, then 3, then bob, amara and zara, gives found at 14, not found and found at 0.
const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];
// your code: read with next() and nextInt(), push every line of output to out
console.log(out.join("\n"));
Not graded on its own. Zara put zara last in the sample on purpose: its index is 0. Check that your program does not call it missing.
Amara writes the club newsletter. It lists who came to the meeting as one English sentence, and some people signed the sheet twice.
Input. n, then n names, one per line.
Output. One line: each name once, in the order first seen. One name prints alone, two are joined by and , and three or more are separated by , with and before the last.
Constraints. 1 <= n <= 1000. Each name has 1 to 30 characters.
Sample. Input 4, then Amara, Bob, Amara and Kenji, gives Amara, Bob and Kenji.
const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];
// your code: read with next() and nextInt(), push every line of output to out
console.log(out.join("\n"));
Not graded on its own. Arrays have includes too, which helps with the repeats. Then decide which part is a join and which is a template.
Kenji's importer now handles the address, but the club's notes column has quotes inside quotes. Finish the parser so that every record of RFC 4180's shape comes out right.
Input. n, then n records, one per line. Fields are separated by commas. A field is unquoted (any characters except a comma or a double quote) or quoted. A quoted field starts with a double quote and ends with the double quote followed by a comma or the end of the record. Inside it any character may appear, a comma included, and two double quotes in a row stand for one. Every record is well formed, spaces belong to the field, and an empty record is one empty field. This problem uses the starter's lines variant, and the last line of the input is never empty.
Output. One line per record: the number of fields, then every field in square brackets, all separated by single spaces. A quoted field is printed without its outer quotes, and each pair of double quotes inside it as one.
Constraints. 1 <= n <= 10000. Each record has 0 to 500 printable ASCII characters (codes 32 to 126).
Sample. Input 3, then Zara,14,Dhaka, "12, Lake Road",Amara,"She said ""hi""" and ,, on three lines, gives 3 [Zara] [14] [Dhaka], 3 [12, Lake Road] [Amara] [She said "hi"] and 3 [] [] [].
const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];
// your code: read with next() and nextInt(), push every line of output to out
// Read n, then n records. Per record: the field count and every field in brackets.
console.log(out.join("\n"));
Graded as csv-line. Start from Example 1. Inside quotes, a quote is either the end of the field or the first of a pair, and the next character tells you which.
Common doubts
Is
splitfaster than my loop?Kenji asks, and it is the wrong first question. Both visit every character once, and this lesson measured nothing. Choose by correctness: a quoted comma makes
splitwrong at any speed.If
substris legacy, why does Node still run it?The standard keeps it in Annex B, the part that lists old features browsers must keep so old pages still work. MDN marks it deprecated. It works, and the chart shows why new code avoids it.
localeCompareorIntl.Collator?For the same language and options they give the same order.
localeCompareis handy for one comparison. For sorting a long list, MDN advises making oneIntl.Collatorand passing itscomparetosort.Could a regular expression parse CSV instead of the loop?
For simple lines, yes. With quoted fields and doubled quotes, the pattern grows hard to read and harder to test. A loop you can trace, or a tested library such as Papa Parse, is easier to trust.
Key takeaways
splitandjoinwhen the separator can never sit inside a field; a character loop with aninQuotesflag when it can.sliceoversubstring, and neversubstr: negatives count from the end and nothing is swapped.includesfor "is it there",indexOffor "where", compared with -1.- A template literal for a line with values,
+for two pieces,joinfor a list. <for codes;localeCompareorIntl.Collatorfor names people read;[...s]oversplit("")when an emoji may appear; a regular expression when the job says "any".- Go deeper: Under the Hood, immutability, ropes and UTF-16 (Pro).
Next, lesson 05 (Pro) opens the string itself: what += really costs, UTF-16 and code points, and the characters a person sees. On the free path, the problems page comes after it, with this module's ten judged problems.
End of lesson 4
Mark it done, and your progress moves with you.
Next: Under the Hood: Immutability, Ropes and UTF-16