Learn JavaScript

Lesson 1 of 9 · Strings and Template Literals

Module 3 · Strings and Template Literals

A String Is a Row of Code Units: The Picture

FreeReading

In this lesson

  • Picture a string as a numbered row of boxes, and read any box with s[i] or s.at(i).
  • Explain why s[0] = "x" changes nothing, and why every string method hands back a new string.
  • Predict when length is not the number of letters a reader sees, for Bangla text and for emoji.

Maria's school prints name badges. The printer gives each name one small box per character, and it asks the program how many boxes to print: name.length. "Maria" gets 5 boxes, which looks right. Her friend Mahi types "মাহি" and gets 4 boxes for what reads as two letters. Zara adds a smiley to her name and gets one box too many. The printer is not broken. It counts something, and this lesson finds out exactly what.

What a string is

A string is a value that holds text: a name, a message, a line of input. Module 1 met it as one of the seven primitive types. Inside the engine, a string is a row of small boxes, one after another.

Each box holds one code unit: a number from 0 to 65535 that stands for a piece of text. The boxes are numbered from 0, and the number of a box is its index. The number of boxes is the string's length.

For English letters, digits and most everyday characters, one character is one box. Here is the badge printer's question, asked of three names.

console.log("Maria".length);
console.log("মাহি".length);
console.log("Zara 😀".length);
5
4
7

"Maria" is five boxes. The other two answers are the badge printer's surprises, and the last section of this lesson explains both. So length counts boxes, and for plain English text a box is a letter.

Three ways to write a string

A string in your code is written between quotes. JavaScript has three kinds, and all three make the same type of value.

const name = "Amara";
console.log("She said \"hi\"");
console.log('She said "hi"');
console.log(`Hello, ${name}! You have ${2 + 3} new messages.`);
console.log(`Line one
Line two`);
console.log(typeof `x`);
She said "hi"
She said "hi"
Hello, Amara! You have 5 new messages.
Line one
Line two
string

Double quotes and single quotes work the same way. To put a double quote inside double quotes, write \"; the backslash starts an escape, a pair of characters that stands for one. \n is a line break and \\ is one backslash.

Backticks make a template literal. Inside one, ${...} runs any expression and puts its value into the text. A line break in the code is a line break in the string, too. Lesson 02 teaches template literals in full.

This track writes double quotes, and backticks when the text needs a value inside it. So the three quotes are three spellings of one type, and you pick the one that makes the text easiest to read.

A string, its length and its boxes

"text"   'text'   `text with ${expression}`

s.length        the number of boxes (code units)
s[i]            the box at index i, as a string of length 1
s.at(i)         the same, and a negative i counts from the end
s.method(...)   a new string (or a number, a boolean, an array); s itself never changes
  • The quotes mark where the text starts and ends. They are not part of the string.
  • length is a property, so it has no parentheses.
  • i is an index from 0 to s.length - 1. Outside that range, the answer is undefined.
  • A method is a function that belongs to a value, called with a dot. Lesson 02 takes every string method one by one.

The boxes and their numbers

Draw the string "Amara" as boxes. The index of each box sits above it, from 0 to 4. Because the first box is box 0, the last box is always length - 1, here 4.

The string "Amara" as five numbered boxes const s = "Amara"; s.length is 5 s[i] 0 1 2 3 4 5 A m a r a undefined s.at(i) -5 -4 -3 -2 -1 The first box is index 0, so the last box is index length - 1, here 4. at(-1) is the last box, at(-2) the one before it. s[-1] is not: it gives undefined. The dashed box does not exist. Reading index 5 gives undefined, not an error.

The row under the boxes is a second way to count, from the end. It works with s.at(i), not with square brackets. Now watch the engine answer six questions about the same string.

The string "Amara" is five boxes, numbered 0 to 4. s[i] and s.at(i) read one box, s.at(-1) counts from the end, a number past the last box gives undefined, and slice(1, 4) copies boxes 1, 2 and 3.

index01234
code unitAmara
hex0041006D006100720061
ExpressionResultWhat happened
s.length5Five boxes, so the length is 5. The indexes run from 0 to 4.
s[0]"A"Index 0 is the first box.
s[4]"a"Index 4 is the last box: length - 1.
s.at(-1)"a"at() also counts from the end: -1 is the last box.
s[5]undefinedThere is no box 5. Reading past the end gives undefined, not an error.
s.slice(1, 4)"mar"slice copies boxes 1, 2 and 3 into a new string. Box 4 is where it stops, and it is left out.

The last step previewed slice, which copies a range of boxes into a new string. Lesson 02 teaches it with every other method.

Reading one box: s[i] and s.at(i)

Square brackets after a string read one box: s[0] is the first. The answer is itself a string, one box long. s.at(i) does the same, and it also takes a negative index, so s.at(-1) is the last box.

const s = "Amara";
console.log(s[0], s[1], s[4]);
console.log(s[s.length - 1], s.at(-1), s.at(-2));
console.log(s[5], s[-1], s.at(5));
console.log(typeof s[0], s[0].length);
A m a
a a r
undefined undefined undefined
string 1

Line 2 reads the last box twice: once by working out its index, once with at(-1). Line 3 asks for boxes that do not exist. There is no error, only undefined, the value Module 1 called "no value yet".

Note s[-1]: square brackets do not count from the end. They look for a box named "-1", find none, and give undefined. So use s.at(-1) for the last box, and s[i] when you already know the index.

A string never changes

You can read a box, but you cannot change one. A string is immutable: once made, it stays exactly as it is. Writing to a box is not an error in a normal program. It just does nothing.

let word = "kenji";
word[0] = "K";
console.log(word);

const upper = word.toUpperCase();
console.log(word, upper);

word = "K" + word.slice(1);
console.log(word);
kenji
kenji KENJI
Kenji

Line 2 tried to change box 0, and the string ignored it. toUpperCase() did not change word either. It made a new string, "KENJI", and returned it, and the program kept it in upper.

The last change worked because it built a new string and stored it in word. That is allowed: word is a let, so it may hold a different string. The old "kenji" was never touched. So let and const decide whether a name can point at a new value, and immutability means the string itself never changes.

Every string method works this way: it returns a new value and leaves the original alone. Strict mode is a stricter setting of the language, and the modules of Module 11 turn it on. In strict mode, the write on line 2 stops the program with a TypeError, which lesson 06 shows.

When length is not the number of letters

Back to the badge printer. A box holds a number from 0 to 65535, and most of the world's letters fit in one box. Two things do not work the way a reader expects.

  • Bangla vowel signs are boxes of their own. মা is the letter ম followed by the vowel sign া. Unicode, the standard that numbers every character, stores the sign as its own character. A reader sees one letter, and the string holds two boxes.
  • An emoji often needs two boxes. Unicode now numbers more characters than fit in 65535. A character past that limit, such as 😀, is stored as two boxes called a surrogate pair. Each half means nothing alone.

The way JavaScript stores text in these boxes is called UTF-16. The widget below shows "মাহি 😀" box by box, with what a reader sees drawn underneath.

The string "মাহি 😀" is seven code units: the four of মাহি (two letters and two vowel signs), a space, and two for the emoji. A reader sees four characters, so length counts boxes, not letters.

index0123456
code unitমাহিspace0xD83D0xDE00
hex09AE09BE09B909BF0020D83DDE00
character a reader sees0 to 12 to 345 to 6
ExpressionResultWhat happened
s.length7Seven boxes. A reader sees four characters: মা, হি, a space and the emoji.
s[0]"ম"Box 0 holds the letter ম.
s[1]"া"Box 1 holds only the vowel sign. In Unicode it is a character of its own, drawn onto the letter before it.
s.slice(0, 2)"মা"Boxes 0 and 1 together make মা, the one letter a reader sees.
s[5]"\ud83d"The emoji needs two boxes. Box 5 is only its first half, which prints as nothing useful on its own.
s.slice(5, 7)"😀"Boxes 5 and 6 together are the emoji.
[...s].length6Spreading the string counts code points, so the emoji's two halves count once: 6. The vowel signs still count on their own.
[...new Intl.Segmenter().segment(s)].length4Counting what a reader sees needs Intl.Segmenter: 4. Lesson 05 opens it up.

Each character has a Unicode number, its code point. The emoji is one code point stored in two boxes. [...s] spreads a string into a list of its code points, and Module 7 teaches that spread form for lists. So [...s].length counts the emoji once.

const s = "মাহি 😀";
console.log(s.length);
console.log([...s].length);
console.log(JSON.stringify(s[5]), JSON.stringify(s.slice(5)));
7
6
"\ud83d" "😀"

JSON.stringify writes a value as code, so the lone half shows up as its number, \ud83d, instead of a broken symbol. Three counts, three answers: 7 boxes, 6 code points, 4 characters a reader sees.

For this track's problems, length is the count the statement means unless it says otherwise. Counting what a reader sees takes a tool from lesson 05. So when a string may hold Bangla or emoji, never call length "the number of letters".

Example 1: the boxes of one word

The smallest program in the lesson. It reads boxes from the front, from the back and past the end.

const word = "Zara";
console.log(word.length);
console.log(word[0], word[3]);
console.log(word.at(-1), word.at(-4));
console.log(word[4]);
4
Z a
a Z
undefined

Four boxes, so the indexes are 0 to 3, and word[4] is past the end. at(-4) counts four from the end and lands on the first box.

Run in Compiler
Example 2: Amara's name card

Amara's program reads one name with the fixed starter and prints a small card about it. A template literal puts each value into its line.

const input = require("fs").readFileSync(0, "utf8");
const tokens = input.split(/\s+/).filter(Boolean);
let at = 0;
const next = () => tokens[at++];
const nextInt = () => Number(next());
const out = [];

const name = next();
out.push(`Name:   ${name}`);
out.push(`Boxes:  ${name.length}`);
out.push(`First:  ${name[0]}`);
out.push(`Last:   ${name.at(-1)}`);
out.push(`Loud:   ${name.toUpperCase()}`);
out.push(`Still:  ${name}`);

console.log(out.join("\n"));
Name:   Amara
Boxes:  5
First:  A
Last:   a
Loud:   AMARA
Still:  Amara

That output is for the input Amara. The last line is the proof of immutability: toUpperCase() made "AMARA" for the line before it, and name is still "Amara".

Run in Compiler
Example 3: the badge printer, honestly

Maria rewrites the badge printer's check. It reads n names and prints, for each, the boxes length counts and the code points [...name] counts. A difference means an emoji, so the badge needs a closer look.

const input = require("fs").readFileSync(0, "utf8");
const tokens = input.split(/\s+/).filter(Boolean);
let at = 0;
const next = () => tokens[at++];
const nextInt = () => Number(next());
const out = [];

const n = nextInt();
for (let i = 0; i < n; i++) {
  const name = next();
  const boxes = name.length;
  const points = [...name].length;
  let line = `${name}: ${boxes} boxes, ${points} code points`;
  if (boxes !== points) {
    line += ", check the badge";
  }
  out.push(line);
}

console.log(out.join("\n"));
Maria: 5 boxes, 5 code points
মাহি: 4 boxes, 4 code points
Zara😀: 6 boxes, 5 code points, check the badge
Kenji: 5 boxes, 5 code points

That output is for the input 4, then Maria মাহি Zara😀 Kenji. The for loop and the if are the ones you met in Module 1; Module 4 teaches them properly. += adds text to the end of line, which builds a new string each time.

The check catches the emoji, and it does not catch মাহি: four code points, two letters a reader sees. Lesson 05 shows the tool that counts those, Intl.Segmenter.

Run in Compiler

Where this is used

  • Web form limits. The HTML standard measures the maxlength and minlength of a text field in UTF-16 code units, the same count as length. A field with maxlength="10" takes ten Latin letters but only five 😀.
  • Databases. In MySQL, a VARCHAR(10) column with the utf8mb4 character set holds ten characters, and 😀 counts as one. A JavaScript check with length counts it as two, so the form and the database can disagree about the same name.
  • Post length on X. X's open-source twitter-text library does not use length to count a post. It gives each character a weight, and an emoji counts the same however many code units it takes.

Common mistakes

1. Reading the last box at index length.

const name = "Bob";
console.log(name[name.length]);
console.log(name[name.length - 1]);
undefined
b

No error, just undefined where a letter should be. The boxes run from 0 to length - 1, so the last one is name[name.length - 1], or name.at(-1). You will make this mistake because you count "the 3rd box" from 1, and the index counts from 0.

2. Expecting a write to a box to change the string.

let title = "hello";
title[0] = "H";
console.log(title);
hello

The program runs, prints hello, and gives no message at all. Build a new string and store it: title = "H" + title.slice(1);. You will try the write because square brackets read a box, and writing looks like the same move.

3. Calling a method and dropping its answer.

let city = "dhaka";
city.toUpperCase();
console.log(city);
dhaka

The method made "DHAKA" and nobody kept it. Store the answer: city = city.toUpperCase();. You will write this because the name reads like an order to the string, and a string takes no orders.

4. Counting from the end with square brackets.

const code = "BD-1207";
console.log(code[-1]);
console.log(code.at(-1));
undefined
7

code[-1] looks for a box named "-1" and finds none. Use at(-1). You will write [-1] if you know Python, where it works.

Brain teaser

Bob is sure all three lines print the last letter of the name.

console.log("Zara"[4]);
console.log("Zara".at(-5));
console.log("Zara".length - 1);

What does each line really print, and which one is closest to what Bob meant?

Count the boxes of "Zara" from 0, then from -1 backwards. The third line never reads a box at all.

Exercise 1Easy

Kenji wants a quick look at any word: its first box, its last box and how many boxes it has.

Input. One word of ASCII letters.

Output. Three lines: the first letter, the last letter, and the length.

Constraints. The word has 1 to 50 letters.

Sample. Input Kenji gives K, i and 5 on three lines.

const input = require("fs").readFileSync(0, "utf8");
const tokens = input.split(/\s+/).filter(Boolean);
let at = 0;
const next = () => tokens[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out

console.log(out.join("\n"));

Not graded on its own. Zara's first test is a word of one letter: what are its first and last boxes?

Run in Compiler
Exercise 2Medium

Bob prints the middle of each word on a game board. A word of odd length has one middle letter. A word of even length has two, and he wants both.

Input. n, then n words of ASCII letters.

Output. One line per word: its middle letter, or its two middle letters.

Constraints. 1 <= n <= 1000. Each word has 1 to 50 letters.

Sample. Input 3 and Amara Zara a gives a, ar and a on three lines.

const input = require("fs").readFileSync(0, "utf8");
const tokens = input.split(/\s+/).filter(Boolean);
let at = 0;
const next = () => tokens[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out

console.log(out.join("\n"));

Not graded on its own. Math.floor(word.length / 2) is the index just right of the middle. Work out the even case on "Zara" by hand first.

Run in Compiler
Exercise 3Medium

Zara's chat app shows how long each message is, and it got Bangla and emoji wrong. She wants both counts side by side for every message.

Input. This problem uses the lines variant of the starter. Line 1 holds n. Then n lines of text follow, in UTF-8. The last line of the input is never empty.

Output. One line per text line: its length in UTF-16 code units, a space, and its number of code points.

Constraints. 1 <= n <= 10000. Each line has 0 to 1000 code points, no carriage return and no lone surrogate.

Sample. Input 3, Maria, মাহি 😀 and 👋🏽 hi on four lines gives 5 5, 7 6 and 7 5.

const input = require("fs").readFileSync(0, "utf8");
const lines = input.split("\n");
let at = 0;
const next = () => lines[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out
// Read n, then n lines. Per line: code units, then code points.

console.log(out.join("\n"));

Graded as true-length. The lines variant keeps spaces and empty lines, so next() returns a whole line. The hidden tests include empty lines, flags and emoji with skin tones.

Run in Compiler
Exercise 4Hard

Zara's game shows the last character of each player's name on a tile. Some names end in an emoji, and at(-1) puts half an emoji on the tile.

Input. n, then n names. A name is ASCII letters, possibly followed by one emoji such as 😀 or 🐱 (one code point each). No name contains a space.

Output. One line per name: its last code point, so a whole emoji when the name ends in one.

Constraints. 1 <= n <= 1000. Each name has 1 to 30 code points.

Sample. Input 3 and Zara😀 Kenji Bob🐱 gives 😀, i and 🐱 on three lines.

const input = require("fs").readFileSync(0, "utf8");
const tokens = input.split(/\s+/).filter(Boolean);
let at = 0;
const next = () => tokens[at++];
const nextInt = () => Number(next());
const out = [];

// your code: read with next() and nextInt(), push every line of output to out

console.log(out.join("\n"));

Not graded on its own. The lesson pointed at the idea: [...name] is a list of code points, and a list has at() too.

Run in Compiler

Common doubts

  • Why does JavaScript count code units and not letters?

    When JavaScript was made in 1995, every Unicode character fit in 16 bits, so one box was one character. Unicode grew past that limit in 1996, and JavaScript, like Java, kept its boxes and stored the new characters as pairs. Changing length now would break old programs. So length counts boxes, and the newer tools count the rest.

  • Is a string an array?

    No, though it looks like one: it has a length and numbered boxes you read with square brackets. A string is a primitive value and can never change; an array is an object you can change, which Module 7 teaches. You will turn one into the other often, with split and join in lesson 02.

  • Single or double quotes: does it matter?

    Not to the engine: both make the same string. Teams pick one and let a formatter keep it consistent; Prettier's default is double quotes, and so is this track's. Pick the other one only to avoid escaping, as in 'She said "hi"'.

  • If a string never changes, how do programs edit text all day?

    They build a new string from pieces of the old one and store it under the same name. "K" + word.slice(1) above is the pattern. Lesson 02 gives you every piece-cutting method, and lesson 05 shows why building new strings is cheaper than it sounds.

  • Does const make a string unchangeable?

    The string is unchangeable anyway, with let or const. const only stops the name from pointing at a different value later. So const word = "hi"; word = "ho"; fails because of const, and word[0] = "H" does nothing because of the string.

Key takeaways

  • A string is a row of boxes (code units), indexed from 0 to length - 1; s[i] and s.at(i) read one box.
  • s.at(-1) is the last box; s[-1] and any index past the end give undefined, not an error.
  • A string is immutable: a write to a box does nothing, and every method returns a new value.
  • Strings take double quotes, single quotes or backticks; backticks make a template literal with ${...} and real line breaks.
  • length counts UTF-16 code units: a Bangla vowel sign is its own box, and an emoji is often two.
  • Go deeper: Under the Hood, immutability, ropes and UTF-16 (Pro).

Next, lesson 02 takes every string method one by one, with its output and its cost, and shows the full table of character codes.

End of lesson 1

Mark it done, and your progress moves with you.

Next: Every String Method, One by One