Learn JavaScript

Lesson 4 of 9 · Strings and Template Literals

Module 3 · Strings and Template Literals

When to Use Which: Methods or a Loop, a Template or +, and When It Is a Regex

FreeReading

In this lesson

  • Choose split and join when the text has a clean separator, and a character loop with an inQuotes flag when it does not.
  • Pick slice over substring and substr, includes or indexOf, a template literal, + or join, and < or localeCompare, and give the reason for each.
  • Use the chart and the flowchart to take text apart, keep an emoji whole with [...s], and recognise when a job has become a regular expression.

Kenji imports the chess club's member list from a spreadsheet. He wrote the importer in ten minutes and made it fast: one split(",") per line. Then the club's treasurer asks why Amara lives in a town called Lake Road". Her address, 12, Lake Road, held a comma, and every column after it moved one place right. Fast and wrong is still wrong. This lesson finds the point where split stops working, then makes six more judgement calls, each with a program.

Where split stops working

A CSV file puts a field in double quotes when the field holds a comma. That is the rule from the CSV standard, RFC 4180. split(",") knows nothing about quotes, so it cuts inside them.

const record = "Kenji,\"12, Lake Road\",Dhaka";
const fields = record.split(",");
console.log(fields.length);
console.log(fields);
4
[ 'Kenji', '"12', ' Lake Road"', 'Dhaka' ]

Three fields went in and four came out, with the quote marks stuck to the pieces. No error appears, so the bug travels into the spreadsheet. split is the right tool only when the separator can never appear inside a field. Here it can, so the text needs a reader that remembers whether it is inside quotes.

A loop that knows it is inside quotes

A flag is a boolean variable that remembers one fact while a loop runs. Here the fact is "am I inside quotes?". A quote mark flips the flag. A comma ends a field only when the flag is off. Every other character joins the current field.

The quote-aware loop

let inQuotes = false;
for (const ch of record) {
  if (ch === "\"")                 inQuotes = !inQuotes;
  else if (ch === "," && !inQuotes) end the field
  else                             add ch to the field
}
  • inQuotes starts false, because a record starts outside quotes.
  • "\"" is a string holding one double quote; the backslash keeps it from closing the string.
  • !inQuotes is true only outside quotes, so a comma inside them is just text.
Example 1: Kenji's record, field by field

The program reads one record and prints its fields one per line, with split's count beside the loop's.

const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];

const record = next();
const fields = [];
let field = "";
let inQuotes = false;
for (const ch of record) {
  if (ch === "\"") {
    inQuotes = !inQuotes;
  } else if (ch === "," && !inQuotes) {
    fields.push(field);
    field = "";
  } else {
    field += ch;
  }
}
fields.push(field);

out.push(`split(","): ${record.split(",").length} pieces`);
out.push(`the loop:   ${fields.length} fields`);
for (let i = 0; i < fields.length; i++) {
  out.push(`${i + 1}: ${fields[i]}`);
}

console.log(out.join("\n"));
split(","): 4 pieces
the loop:   3 fields
1: Kenji
2: 12, Lake Road
3: Dhaka

That output is for the input Kenji,"12, Lake Road",Dhaka. The comma inside the quotes stayed in the address, and the quote marks themselves were dropped. The last push after the loop saves the final field, which no comma ends.

One honest gap remains. RFC 4180 writes a quote inside a quoted field as two quotes, "". This loop flips the flag twice and loses both, so "She said ""hi""" comes out as She said hi. Exercise 3 asks you to keep it.

Run in Compiler

Here is the whole decision as one picture. Start at the top, answer each question, and stop at the first yes.

How do I take this text apart? A decision flowchart How do I take this text apart? Can a separator hide inside quotes or after an escape? yes a character loop with an inQuotes flag no Is the separator one fixed text, such as a comma, a space or a tab? yes split(sep) and join(sep) to rebuild no Is it a pattern: any run of spaces, any digit, any capital letter? yes a regular expression (Module 16) no Do you need the characters a person sees (Bangla, a waving hand)? yes Intl.Segmenter (lesson 05) no code points: [...s] or Array.from(s) never split("") if an emoji may appear Ask the questions in order. A quoted field breaks split even when the separator is fixed.

slice, substring and substr: the chart

Lesson 02 met three ways to cut a piece out of a string. They agree on the easy calls and disagree on the rest. The chart runs all three on the same six calls, so the difference is visible.

slice, substring and substr on six calls s = "template": six calls, three methods arguments s.slice s.substring s.substr (1, 4) "emp" "emp" "empl" (4, 1) "" "emp" "l" (-3) "ate" "template" "ate" (2, -1) "mplat" "te" "" (-4, 2) "" "te" "la" (3) "plate" "plate" "plate" negative index counts from the end becomes 0 start only start after end gives "" swaps them 2nd is a length status use it works, surprises legacy (Annex B) All three agree only on (3). substring hides a swapped pair; substr's 4 in (1, 4) is a length. slice is the one whose answer you can predict from the indexes alone.
Example 2: the chart, run

Each line prints the three answers for one row, in brackets, so an empty string still shows.

const s = "template";
console.log(`(1, 4):  [${s.slice(1, 4)}] [${s.substring(1, 4)}] [${s.substr(1, 4)}]`);
console.log(`(4, 1):  [${s.slice(4, 1)}] [${s.substring(4, 1)}] [${s.substr(4, 1)}]`);
console.log(`(-3):    [${s.slice(-3)}] [${s.substring(-3)}] [${s.substr(-3)}]`);
console.log(`(2, -1): [${s.slice(2, -1)}] [${s.substring(2, -1)}] [${s.substr(2, -1)}]`);
console.log(`(-4, 2): [${s.slice(-4, 2)}] [${s.substring(-4, 2)}] [${s.substr(-4, 2)}]`);
console.log(`(3):     [${s.slice(3)}] [${s.substring(3)}] [${s.substr(3)}]`);
(1, 4):  [emp] [emp] [empl]
(4, 1):  [] [emp] [l]
(-3):    [ate] [template] [ate]
(2, -1): [mplat] [te] []
(-4, 2): [] [te] [la]
(3):     [plate] [plate] [plate]

Row (4, 1) is Bob's off-by-one waiting to happen. slice says "nothing there" and the empty result shows up in a test. substring quietly swaps the two numbers, and the wrong piece looks fine. So write slice: negatives count from the end, the end is excluded, and nothing is swapped behind your back.

Run in Compiler

includes for "is it there", indexOf for "where"

Both search a string. includes answers a yes-or-no question with a boolean. indexOf answers "where?" with a number, and -1 for "nowhere". Pick by the question you are asking.

Example 3: an email address

Is there an @? Where is it? And what is the name in front of it?

const email = "kenji@mail.com";
console.log(email.includes("@"));
console.log(email.indexOf("@"));
console.log(email.slice(0, email.indexOf("@")));
console.log(email.includes("#"), email.indexOf("#"));
true
5
kenji
false -1

For "is it there", includes reads like the question and cannot fall into lesson 02's trap, where if (s.indexOf(x)) treats a match at index 0 as false. When you need the position, as for the name, indexOf is the tool, and you compare its answer with -1.

Run in Compiler

A template for a line with values, + for two pieces, join for many

Three ways build text, and each has its job. A template literal shows the finished line with holes in it, so you read it as it will print. + is fine for two pieces. join glues a whole array with one separator.

Example 4: one message, three builders

The same sentence with + and as a template, then a short label, then a CSV row from an array.

const name = "Amara";
const books = 3;
const total = 1940;
console.log("Dear " + name + ", your " + books + " books cost Tk " + total + ".");
console.log(`Dear ${name}, your ${books} books cost Tk ${total}.`);
console.log("Tk " + total);
const row = ["Amara", "16", "Dhaka", "amara@mail.com"];
console.log(row.join(","));
Dear Amara, your 3 books cost Tk 1940.
Dear Amara, your 3 books cost Tk 1940.
Tk 1940
Amara,16,Dhaka,amara@mail.com

The first two lines print the same text. In the + line, count the quote marks and the spaces you had to place by hand; one missing space and the receipt says 3books. For two pieces, "Tk " + total is short and clear. For a list, join puts exactly one separator between items and none at the ends.

Many teams make the template a rule. ESLint's prefer-template rule flags + between text and a value, and it can rewrite the line for you. That is a style choice that a team agrees on, not a speed rule: both lines produce the same string.

Run in Compiler

< for codes, localeCompare or Intl.Collator for people

Lesson 02 showed that < compares strings code by code, so every capital comes before every small letter. That is right for codes and IDs. It is wrong for a list of names a person will read. localeCompare compares the way a language orders words. An Intl.Collator is an object that does the same comparison with its settings fixed once, which suits a long sort.

Example 5: the club's sign-up sheet, sorted

Members typed their own names, some with a capital and one with an accent. The default sort() compares as text in code order. The arrow function inside sort is the short form Module 5 explains.

const names = ["zara", "Amara", "Émile", "bob", "Kenji"];
names.sort();
console.log(`< order:        ${names.join(", ")}`);
names.sort((x, y) => x.localeCompare(y));
console.log(`localeCompare:  ${names.join(", ")}`);
const collator = new Intl.Collator("en");
names.sort(collator.compare);
console.log(`Intl.Collator:  ${names.join(", ")}`);
const photos = ["photo10.jpg", "photo9.jpg", "photo2.jpg"];
photos.sort(new Intl.Collator("en", { numeric: true }).compare);
console.log(`numeric: true:  ${photos.join(", ")}`);
< order:        Amara, Kenji, bob, zara, Émile
localeCompare:  Amara, bob, Émile, Kenji, zara
Intl.Collator:  Amara, bob, Émile, Kenji, zara
numeric: true:  photo2.jpg, photo9.jpg, photo10.jpg

In code order, bob waits behind Kenji, and Émile lands last, because É is code 201. localeCompare and the collator give the order a reader expects. The option numeric: true compares runs of digits as numbers, so photo10 comes after photo9. For Bangla names, new Intl.Collator("bn") uses the Bangla rules.

Run in Compiler

[...s] over split("") when an emoji may appear

split("") cuts between code units, and lesson 01 showed that an emoji takes two. The spread [...s] and Array.from(s) walk the string by code point instead, so an emoji stays whole. Module 7 shows spread on arrays.

Example 6: Zara reverses her name

JSON.stringify prints each piece in quotes and writes a lone half of an emoji as an escape such as \ud83d. Printed directly, a lone half shows as a replacement character instead.

const name = "Zara 😀";
console.log(name.length);
console.log(JSON.stringify(name.split("")));
console.log(JSON.stringify([...name]));
console.log(JSON.stringify(Array.from(name)));
console.log(JSON.stringify(name.split("").reverse().join("")));
console.log([...name].reverse().join(""));
7
["Z","a","r","a"," ","\ud83d","\ude00"]
["Z","a","r","a"," ","😀"]
["Z","a","r","a"," ","😀"]
"\ude00\ud83d araZ"
😀 araZ

split("") gave seven pieces, the last two half an emoji each, and reversing put the halves in the wrong order. [...name] gave six pieces and a clean reverse. For text that is ASCII by its rules, such as the judge's words, split("") is fine. For anything a person typed, use [...s]. Bangla vowel signs still split from their letters, and that is lesson 05's Intl.Segmenter.

Run in Compiler

When the job has become a regular expression

Every tool above looks for fixed text. Some jobs look for a pattern: any run of spaces, any digit, a line made only of digits. That is the job of a regular expression, a small language for describing text, which Module 16 teaches. Here are two one-liners, named and not taught.

console.log("too   many    spaces".replace(/\s+/g, " "));
console.log(/^\d+$/.test("2026"), /^\d+$/.test("20x6"));
too many spaces
true false

The first squeezes every run of spaces into one. The second asks whether a text is digits from start to end. Each would take a loop of several lines. The sign is the word "any": when you hear yourself say "any run of" or "any digit", the job has become a pattern.

Where this is used

  • ESLint. The rule prefer-template ("Require template literals instead of string concatenation") flags "Hello, " + name + "!". It is not in ESLint's recommended set, so a team turns it on by choice, and --fix rewrites the line.
  • The Airbnb JavaScript Style Guide. Its rule 6.3 says to use template strings instead of concatenation when building up strings, and names prefer-template and template-curly-spacing as the ESLint rules.
  • RFC 4180. The 2005 CSV standard. Its rule 6 puts a field holding a comma in double quotes, which Example 1 handles. Its rule 7 writes a quote inside a field as two quotes, which Exercise 3 handles.
  • ICU. Node and Chrome build Intl.Collator on ICU, the International Components for Unicode library, which follows the Unicode Collation Algorithm. Android hands the same library to apps as android.icu.text.Collator.

Common mistakes

1. Reading substr's second number as an end.

console.log("hello".substr(1, 3), "hello".slice(1, 3));
ell el

substr(1, 3) means three characters from index 1, while slice(1, 3) stops before index 3. No error appears, just one character too many. Write slice. You will meet substr in old code and assume it matches its neighbours.

2. Using includes where a position is needed.

const mail = "kenji@mail.com";
const at = mail.includes("@");
console.log(at);
console.log(mail.slice(at));
true
enji@mail.com

slice wanted a number, got true, and converted it to 1, as Module 2's coercion rules say. The program cut one letter and kept going. Use indexOf when you need where. You will reach for includes because it is the newer, friendlier name.

3. Adding numbers inside a + chain.

console.log("Total: " + 2 + 3);
console.log(`Total: ${2 + 3}`);
Total: 23
Total: 5

+ works left to right, so the text meets 2 first and every + after it joins text. Inside ${}, the sum is worked out on its own. You will write the first line because it worked a moment ago with one number.

Brain teaser

console.log(["b", "a", "B", "A"].sort());
console.log(["b", "a", "B", "A"].sort((x, y) => x.localeCompare(y)));

The two lines sort the same four letters and print different orders. Predict both, and say why they differ. Then decide which of the two a phone's contact list uses.

For the first line, look up the four codes in lesson 02's table. For the second, ask which letters a person expects to find next to each other, and which of a pair might come first.

Exercise 1Easy

Kenji's club keeps every member's email on one line, separated by commas. The secretary types a name to look for, and wants to know whether it is there and where.

Input. Line 1 is the list. Line 2 is n. Then n lines, each a text to look for.

Output. One line per text: found at and the index of its first match, or not found.

Constraints. 1 <= n <= 1000. The list has 1 to 2000 characters; each text has 1 to 50.

Sample. The list zara@mail.com,bob@mail.com,kenji@mail.com, then 3, then bob, amara and zara, gives found at 14, not found and found at 0.

const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out

console.log(out.join("\n"));

Not graded on its own. Zara put zara last in the sample on purpose: its index is 0. Check that your program does not call it missing.

Run in Compiler
Exercise 2Medium

Amara writes the club newsletter. It lists who came to the meeting as one English sentence, and some people signed the sheet twice.

Input. n, then n names, one per line.

Output. One line: each name once, in the order first seen. One name prints alone, two are joined by and , and three or more are separated by , with and before the last.

Constraints. 1 <= n <= 1000. Each name has 1 to 30 characters.

Sample. Input 4, then Amara, Bob, Amara and Kenji, gives Amara, Bob and Kenji.

const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out

console.log(out.join("\n"));

Not graded on its own. Arrays have includes too, which helps with the repeats. Then decide which part is a join and which is a template.

Run in Compiler
Exercise 3Hard

Kenji's importer now handles the address, but the club's notes column has quotes inside quotes. Finish the parser so that every record of RFC 4180's shape comes out right.

Input. n, then n records, one per line. Fields are separated by commas. A field is unquoted (any characters except a comma or a double quote) or quoted. A quoted field starts with a double quote and ends with the double quote followed by a comma or the end of the record. Inside it any character may appear, a comma included, and two double quotes in a row stand for one. Every record is well formed, spaces belong to the field, and an empty record is one empty field. This problem uses the starter's lines variant, and the last line of the input is never empty.

Output. One line per record: the number of fields, then every field in square brackets, all separated by single spaces. A quoted field is printed without its outer quotes, and each pair of double quotes inside it as one.

Constraints. 1 <= n <= 10000. Each record has 0 to 500 printable ASCII characters (codes 32 to 126).

Sample. Input 3, then Zara,14,Dhaka, "12, Lake Road",Amara,"She said ""hi""" and ,, on three lines, gives 3 [Zara] [14] [Dhaka], 3 [12, Lake Road] [Amara] [She said "hi"] and 3 [] [] [].

const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out
// Read n, then n records. Per record: the field count and every field in brackets.

console.log(out.join("\n"));

Graded as csv-line. Start from Example 1. Inside quotes, a quote is either the end of the field or the first of a pair, and the next character tells you which.

Run in Compiler

Common doubts

  • Is split faster than my loop?

    Kenji asks, and it is the wrong first question. Both visit every character once, and this lesson measured nothing. Choose by correctness: a quoted comma makes split wrong at any speed.

  • If substr is legacy, why does Node still run it?

    The standard keeps it in Annex B, the part that lists old features browsers must keep so old pages still work. MDN marks it deprecated. It works, and the chart shows why new code avoids it.

  • localeCompare or Intl.Collator?

    For the same language and options they give the same order. localeCompare is handy for one comparison. For sorting a long list, MDN advises making one Intl.Collator and passing its compare to sort.

  • Could a regular expression parse CSV instead of the loop?

    For simple lines, yes. With quoted fields and doubled quotes, the pattern grows hard to read and harder to test. A loop you can trace, or a tested library such as Papa Parse, is easier to trust.

Key takeaways

  • split and join when the separator can never sit inside a field; a character loop with an inQuotes flag when it can.
  • slice over substring, and never substr: negatives count from the end and nothing is swapped.
  • includes for "is it there", indexOf for "where", compared with -1.
  • A template literal for a line with values, + for two pieces, join for a list.
  • < for codes; localeCompare or Intl.Collator for names people read; [...s] over split("") when an emoji may appear; a regular expression when the job says "any".
  • Go deeper: Under the Hood, immutability, ropes and UTF-16 (Pro).

Next, lesson 05 (Pro) opens the string itself: what += really costs, UTF-16 and code points, and the characters a person sees. On the free path, the problems page comes after it, with this module's ten judged problems.

End of lesson 4

Mark it done, and your progress moves with you.

Next: Under the Hood: Immutability, Ropes and UTF-16

When to Use Which: Methods or a Loop, a Template or +, and When It Is a Regex | Learn JavaScript | Progsity