Module 3 · string: Text That Knows Its Own Length
Full Programs: From One Word to a Text Report
In this lesson
- Write six complete string programs, from counting words to a marks tool that reads a CSV file with names that contain spaces.
- Split a line at a character in two ways, with
findandsubstrand withgetline(in, field, ','), and say when each one fits. - Use string streams both ways:
istringstreamto read values out of a line,ostringstreamto build a report before printing it.
In C you wrote a marks report with char name[20] for every student. A name with a space broke scanf("%s"), and a name of twenty letters broke the array. This lesson writes that report again with strings, and the name can be as long as the person's real name. Six programs get there, each a little bigger than the last. Read the "new thing" line above each program first; it is what that program is for.
Program 1: count the words and their characters
The new thing: while (cin >> word) reads one word per turn until the input ends, and size() measures each word as it arrives.
A word here is a run of characters with no space in it, which is exactly what >> reads into a string. Lesson 01 showed one read; this program loops. Zara tests the empty input first, so it gets its own line.
#include <iostream>
#include <string>
using namespace std;
int main() {
ios::sync_with_stdio(false);
cin.tie(nullptr);
string word;
int count = 0;
int characters = 0;
int long_words = 0;
while (cin >> word) {
count++;
characters += word.size();
if (word.size() >= 5) {
long_words++;
}
}
if (count == 0) {
cout << "no words\n";
return 0;
}
cout << "words: " << count << '\n';
cout << "characters: " << characters << '\n';
cout << "5 or more characters: " << long_words << '\n';
return 0;
}
words: 9
characters: 35
5 or more characters: 3
That output is for the input the quick brown fox and jumps over the lazy dog on two lines. The newline between the lines is just more space to >>. Each word arrives in a string sized to fit it, and size() gives its length at once: quick, brown and jumps have five. So the loop reads any number of words of any length, with no size you chose.
Program 2: a letter count with vector<int>(26)
The new thing: vector<int> count(26, 0) with tolower(u) - 'a' as the index, where u is the character cast to unsigned char.
Amara wants to know if a sentence is a pangram, a sentence that uses all 26 letters. Each letter gets a box: 'a' is 97 and 'z' is 122, so c - 'a' runs from 0 to 25. isalpha and tolower come from <cctype>; the first asks "is this a letter?", the second turns a capital into a small letter.
#include <cctype>
#include <iostream>
#include <string>
#include <vector>
using namespace std;
int main() {
string line;
while (getline(cin, line)) {
vector<int> count(26, 0);
for (char c : line) {
unsigned char u = c;
if (isalpha(u)) {
count[tolower(u) - 'a']++;
}
}
string missing;
int top = 0;
for (int i = 0; i < 26; i++) {
if (count[i] == 0) {
missing += char('a' + i);
}
if (count[i] > count[top]) {
top = i;
}
}
cout << line << '\n';
cout << " most used: " << char('a' + top) << ", " << count[top] << " times\n";
if (missing.empty()) {
cout << " a pangram: all 26 letters\n";
} else {
cout << " missing " << missing.size() << ": " << missing << '\n';
}
}
return 0;
}
Sphinx of black quartz, judge my vow
most used: a, 2 times
a pangram: all 26 letters
Kenji fixed one bug and found two more
most used: e, 4 times
missing 9: chlpqsvyz
Why the cast? On GCC for a PC, char is signed, so a byte above 127 becomes a negative number. Each byte of a Bangla letter is such a byte. The <cctype> functions are defined only for values an unsigned char can hold, so a negative one is undefined behaviour. Copying c into unsigned char u first makes every byte 0 to 255, and isalpha then says no to those bytes. So the vector is the table, and the cast keeps the table safe from text that is not English.
Program 3: split a line at its commas
The new thing: two ways to cut a line into fields: find with substr, and getline(in, field, ',') on an istringstream.
Each line of Amara's class file is a name and three marks with commas between them. The first way does the cutting by hand. line.find(',', start) gives the position of the next comma at or after start, or string::npos when there is none. A field runs from start to just before that comma.
#include <iostream>
#include <string>
using namespace std;
int main() {
string line;
getline(cin, line);
size_t start = 0;
int field = 1;
while (true) {
size_t comma = line.find(',', start);
if (comma == string::npos) {
cout << field << ": [" << line.substr(start) << "] to the end\n";
break;
}
cout << field << ": [" << line.substr(start, comma - start) << "] comma at " << comma << '\n';
start = comma + 1;
field++;
}
return 0;
}
1: [Amara Okafor] comma at 12
2: [85] comma at 15
3: [90] comma at 18
4: [77] to the end
That output is for the input line Amara Okafor,85,90,77. The second argument of substr is a length, not an end, so it is comma - start. The brackets show that the space inside the name stayed, which >> would have lost.
The second way lets a stream do the cutting. Lesson 02 showed getline with a third argument: it reads up to that character and drops it. Put the line in an istringstream, and each call hands you the next field.
#include <iostream>
#include <sstream>
#include <string>
using namespace std;
int main() {
string line;
getline(cin, line);
istringstream in(line);
string field;
int k = 1;
while (getline(in, field, ',')) {
cout << k << ": [" << field << "]\n";
k++;
}
return 0;
}
1: [Amara Okafor]
2: [85]
3: [90]
4: [77]
The same four fields, from the same input, with no position to keep track of. Here is how the two compare.
| Question | find and substr | getline(in, field, ',') |
|---|---|---|
| Do you get each field's position? | yes, find returns it | no, only the text |
Can the separator be longer than one character, like ", "? | yes, find takes a string too | no, it is one char |
| How much code? | a loop, a start and a length to get right | one loop condition |
| What does it need? | <string> only | <sstream> and a copy of the line in the stream |
So reach for getline when you only want the fields, and for find when you need positions or a longer separator.
Program 4: numbers out of each line
The new thing: a new istringstream for every line, read with >> until it runs dry, so an empty line simply reads no numbers.
Alice's phone logs her walks. Each line is one day, with the steps of each walk on it, and a day with no walk is an empty line. She wants each day's walks and steps, and the total of all days.
#include <iostream>
#include <sstream>
#include <string>
using namespace std;
int main() {
ios::sync_with_stdio(false);
cin.tie(nullptr);
string line;
int day = 0;
long long all_steps = 0;
while (getline(cin, line)) {
day++;
istringstream in(line);
int steps;
int walks = 0;
long long total = 0;
while (in >> steps) {
walks++;
total += steps;
}
cout << "day " << day << ": ";
if (walks == 0) {
cout << "no walks\n";
} else {
cout << walks << (walks == 1 ? " walk, " : " walks, ") << total << " steps\n";
}
all_steps += total;
}
cout << "all " << day << " days: " << all_steps << " steps\n";
return 0;
}
day 1: 3 walks, 5400 steps
day 2: no walks
day 3: 1 walk, 5000 steps
day 4: 2 walks, 5000 steps
all 4 days: 15400 steps
That output is for four input lines: 1200 3400 800, an empty line, 5000 and 2500 2500. The outer loop reads lines, so it sees the empty one; cin >> alone would skip it and lose day 2. The stream is made inside the loop, so each day starts fresh. So the line keeps the days apart, and the string stream reads the numbers inside one day.
Program 5: a receipt built in an ostringstream
The new thing: the whole body of a report written into an ostringstream with setw and setprecision, then printed once with str().
David's shop prints a receipt. Its first line says how many lines it has, a number known only after the last item is read. So the program writes the item lines into a string first, prints the heading, then prints the string. setw(n) pads the next value to n characters; left and right choose the side; fixed with setprecision(2) gives two digits after the point. All of them come from <iomanip> and work on any output stream.
#include <iomanip>
#include <iostream>
#include <sstream>
#include <string>
using namespace std;
int main() {
ostringstream body;
body << fixed << setprecision(2);
string item;
int qty;
double price;
double total = 0;
int lines = 0;
while (cin >> item >> qty >> price) {
double cost = qty * price;
body << left << setw(10) << item << right << setw(4) << qty
<< " x " << setw(6) << price << " = " << setw(8) << cost << '\n';
total += cost;
lines++;
}
body << string(34, '-') << '\n';
body << left << setw(26) << "TOTAL" << right << setw(8) << total << '\n';
cout << "David's shop, " << lines << " lines\n";
cout << body.str();
return 0;
}
David's shop, 4 lines
pencil 3 x 0.50 = 1.50
notebook 2 x 2.25 = 4.50
eraser 1 x 0.80 = 0.80
ruler 1 x 1.20 = 1.20
----------------------------------
TOTAL 8.00
That output is for four input lines: pencil 3 0.50, notebook 2 2.25, eraser 1 0.80 and ruler 1 1.20. Each item line is 10 + 4 + 3 + 6 + 3 + 8 = 34 characters, so the dashes are string(34, '-'). The settings stay on the stream, so fixed is written once. So an ostringstream is a draft you can finish before anyone sees it.
Program 6: the CSV marks toolIntermediate
The new thing: the pieces together in a real tool: getline for lines, getline(in, field, ',') for fields, stoi for marks, and a column as wide as the longest name.
A CSV file, for "comma-separated values", stores a table as text: one row per line, commas between the fields. Every spreadsheet program can save a sheet as CSV, so marks lists often travel that way. Each line here is name,mark,mark,mark. The tool prints a table with each student's average, then each test's class average. stoi, from lesson 02, turns a field like "85" into the int 85.
#include <algorithm>
#include <iomanip>
#include <iostream>
#include <sstream>
#include <string>
#include <utility>
#include <vector>
using namespace std;
int main() {
vector<pair<string, vector<int>>> rows;
size_t width = 4; // at least as wide as the word "name"
string line;
while (getline(cin, line)) {
if (line.empty()) {
continue;
}
istringstream in(line);
string name, field;
getline(in, name, ',');
vector<int> marks;
while (getline(in, field, ',')) {
marks.push_back(stoi(field));
}
width = max(width, name.size());
rows.push_back({name, marks});
}
ostringstream table;
table << fixed << setprecision(1);
table << left << setw(width) << "name" << right
<< setw(5) << "T1" << setw(5) << "T2" << setw(5) << "T3" << setw(9) << "average" << '\n';
vector<int> test_sum(3, 0);
for (const auto& [name, marks] : rows) {
int total = 0;
table << left << setw(width) << name << right;
for (int t = 0; t < 3; t++) {
table << setw(5) << marks[t];
total += marks[t];
test_sum[t] += marks[t];
}
table << setw(9) << total / 3.0 << '\n';
}
table << left << setw(width) << "class" << right;
for (int t = 0; t < 3; t++) {
table << setw(5) << (double)test_sum[t] / rows.size();
}
table << '\n';
cout << rows.size() << " students\n" << table.str();
return 0;
}
5 students
name T1 T2 T3 average
Amara Okafor 85 90 77 84.0
Bob Smith 48 55 61 54.7
Maria Lopez Garcia 90 94 88 90.7
Kenji Sato 81 79 85 81.7
Zara Ali 67 72 70 69.7
class 74.2 78.0 76.2
That output is for five input lines: Amara Okafor,85,90,77, Bob Smith,48,55,61, Maria Lopez Garcia,90,94,88, Kenji Sato,81,79,85 and Zara Ali,67,72,70.
The name column is 18 wide because the longest name, Maria Lopez Garcia, has 18 characters. Nobody chose that number; the data did. In C, char name[20] would have cut a longer name, and here a name of 200 characters only makes the column wider. total / 3.0 divides by a double, so 164 / 3 prints as 54.7, not 54. So the C report's two limits, the name length and the space inside a name, are both gone.
Maria signs her notes with initials. An istringstream hands out the words of the name, and the first character of each one is kept, followed by a dot.
#include <iostream>
#include <sstream>
#include <string>
using namespace std;
string initials(const string& full_name) {
istringstream in(full_name);
string word, result;
while (in >> word) {
result += word[0];
result += '.';
}
return result;
}
int main() {
cout << initials("Maria Lopez Garcia") << '\n';
cout << initials(" Amara Okafor ") << '\n';
cout << "[" << initials("") << "]\n";
return 0;
}
M.L.G.
A.O.
[]
Extra spaces cost nothing, because >> skips them. word[0] is safe here, since >> never gives back an empty word.
Alice's counter hands out tickets like A-007. setfill('0') makes setw pad with zeros instead of spaces, and the ostringstream turns the result into a string that can be stored or joined.
#include <iomanip>
#include <iostream>
#include <sstream>
#include <string>
using namespace std;
string ticket_code(char counter, int number) {
ostringstream out;
out << counter << '-' << setw(3) << setfill('0') << number;
return out.str();
}
int main() {
cout << ticket_code('A', 7) << '\n';
cout << ticket_code('B', 42) << '\n';
cout << ticket_code('C', 123) << '\n';
string both = ticket_code('A', 7) + ", " + ticket_code('A', 8);
cout << both << " (" << both.size() << " characters)\n";
return 0;
}
A-007
B-042
C-123
A-007, A-008 (12 characters)
to_string(7) would give "7" with no way to pad it. The stream gives you every <iomanip> setting, and str() gives the text back.
Bob hides messages by moving every letter 3 places along the alphabet, so a becomes d and z wraps round to c. The letter-to-index trick of Program 2 does the work, and % 26 does the wrapping.
#include <cctype>
#include <iostream>
#include <string>
using namespace std;
string shift(string text, int k) {
for (char& c : text) {
unsigned char u = c;
if (islower(u)) {
c = 'a' + (c - 'a' + k) % 26;
} else if (isupper(u)) {
c = 'A' + (c - 'A' + k) % 26;
}
}
return text;
}
int main() {
string secret = shift("Meet Zara at noon", 3);
cout << secret << '\n';
cout << shift(secret, 23) << '\n';
return 0;
}
Phhw Cdud dw qrrq
Meet Zara at noon
Shifting by 23 more makes 26, a full turn, so the message comes back. shift takes its text by value on purpose: it changes the copy and returns it, and the caller's string is untouched.
Where this is used
- Git, the C contrast. Git is written in C, so it wrote its own growable string,
struct strbuf: a buffer, its lengthlenand its roomalloc, always ended by a'\0'. That is whatstd::stringhands you for free. - Linux's /proc files. Files like
/proc/meminforeport memory as lines of text, a name, a number and a unit. Thefreecommand reads that file and splits each line, Program 4's job on a real system. - LLVM's
StringRef. The compiler project's string type has asplitthat takes a separator character and returns the part before its first match and the rest. It is Program 3'sfindandsubstrin one call, on a view of the text that lesson 04 meets asstring_view.
Common mistakes
1. Giving substr an end instead of a length.
size_t comma = line.find(',', start);
cout << "[" << line.substr(start, comma) << "]\n";
No message at any command line. On Amara Okafor,85,90,77 the loop printed [Amara Okafor], then [85,90,77] and [90,77]. The first field was right only because it starts at 0, where an end and a length are the same number. Write substr(start, comma - start). Bob makes this mistake because other languages' slices take an end.
2. One istringstream for every line.
istringstream in;
while (getline(cin, line)) {
in.str(line);
int x, total = 0;
while (in >> x) {
total += x;
}
cout << total << '\n';
}
No message. For the lines 1200 3400 800, 5000 and 2500 2500 it printed 5400, then 0 and 0. The first line's last read failed, and str() gives the stream new text but does not clear that failure. Make the stream inside the loop, as Program 4 does. You will reuse it because making a new object every line feels wasteful.
3. Reading n with >>, then the first row with getline.
cin >> n;
for (int i = 0; i < n; i++) {
getline(cin, line);
istringstream in(line);
getline(in, name, ',');
getline(in, field, ',');
cout << name << ' ' << stoi(field) << '\n';
}
No message at compile time, even with -Wall -Wextra. With the input 2 and two rows, the run stopped with:
terminate called after throwing an instance of 'std::invalid_argument'
what(): stoi
>> left the newline after 2, so the first getline read an empty line, and stoi("") has no number to convert. Finish the line with one more getline right after cin >> n, as Module 1's fast input lesson showed. You will forget it because the input looks like one value per line.
Maria has a text and wants to know how often each letter appears in it. Capital and small letters count as the same letter.
Input. A text of one or more lines, until the input ends.
Output. For each letter that appears, from a to z, one line: the letter, a space, its count. If no letter appears, the word none.
Constraints. The text is at most 1000000 bytes of ASCII.
Sample. Input Hello, World! and Zara 2026 gives nine lines: a 2, d 1, e 1, h 1, l 3, o 2, r 2, w 1 and z 1.
#include <iostream>
#include <string>
#include <vector>
using namespace std;
int main()
{
ios::sync_with_stdio(false);
cin.tie(nullptr);
vector<int> count(26, 0);
string line;
while (getline(cin, line)) {
// For every letter of the line, turn a capital into a small
// letter, then add 1 to count[letter - 'a'].
}
// For each letter from a to z that appears, print the letter,
// a space and its count. Print "none" if no letter appears.
return 0;
}
Graded as letter-frequency, a free problem in this module's problem set.
Each line of Amara's class file holds a student's name and three marks, separated by commas. Names can have spaces inside. She wants the student with the best total.
Input. n, then n lines name,mark1,mark2,mark3.
Output. The best name on line 1 and its total on line 2. On a tie, the first of them.
Constraints. 1 <= n <= 20000. A name has 1 to 40 characters: letters, with single spaces inside, no comma and no space at either end. Each mark is from 0 to 100.
Sample. Input 4, Maria Rose,90,85,77, Bob,100,60,95, Zara Bell,88,95,72 and Kenji,70,70,70 gives Bob and 255 on two lines.
#include <iostream>
#include <sstream>
#include <string>
using namespace std;
int main()
{
ios::sync_with_stdio(false);
cin.tie(nullptr);
int n = 0;
cin >> n;
string line;
getline(cin, line); // finish the line that held n
for (int i = 0; i < n; i++) {
getline(cin, line);
// Split the line at its commas: a name, then three marks.
// Keep the name with the highest total; the first one wins a tie.
}
// Print the best name on one line and its total on the next.
return 0;
}
Graded as csv-best-student, a Pro problem in this module's problem set.
Zara reads every line backwards, word by word. Spaces in the input are messy, but her output must be tidy.
Input. n, then n lines, each with 1 to 1000 words of letters and digits separated by one or more spaces.
Output. Each line's words in reverse order, separated by single spaces.
Constraints. The whole input is at most 1000000 bytes.
Sample. Input 3, the quick brown fox, Zara tests edge cases and hello gives fox brown quick the, cases edge tests Zara and hello on three lines.
#include <iostream>
#include <sstream>
#include <string>
#include <vector>
using namespace std;
int main()
{
ios::sync_with_stdio(false);
cin.tie(nullptr);
int n = 0;
cin >> n;
string line;
getline(cin, line); // finish the line that held n
for (int i = 0; i < n; i++) {
getline(cin, line);
// Read the words of the line, then print them from the last
// to the first, separated by single spaces.
}
return 0;
}
Graded as reverse-words, a Pro problem in this module's problem set.
Real CSV files put a field in double quotes when it holds a comma, so "Okafor, Jr" stays one field. Neither split in Program 3 knows that. Write one that does.
Input. Lines until the end of the input. Fields are separated by commas. A field may be wrapped in double quotes, and then it may contain commas. No quote appears inside a quoted field.
Output. For each line, its fields without the quotes, separated by " | ".
Constraints. At most 1000 lines, each at most 1000 characters, ASCII only.
Sample. Input Amara,"Okafor, Jr",85 and "Hill Road, 12",David,"red, green" gives Amara | Okafor, Jr | 85 and Hill Road, 12 | David | red, green.
#include <iostream>
#include <string>
#include <vector>
using namespace std;
int main() {
ios::sync_with_stdio(false);
cin.tie(nullptr);
string line;
while (getline(cin, line)) {
vector<string> fields;
// Walk the line one character at a time, with a flag that says
// whether you are inside quotes. A comma inside quotes belongs to
// the field; a comma outside quotes ends it. Drop the quotes.
// Print the fields of the line separated by " | ".
}
return 0;
}
Not graded on its own. Neither find nor getline can do it alone: a comma's meaning depends on what came before it, so the loop must remember a state.
Common doubts
Which split should I use in a contest?
Usually
getline(in, field, ','): it is short and hard to get wrong. Usefindandsubstrwhen the separator is longer than one character, or when you need where each field starts.What does
stoido with" 85"or"85 "?Both give 85.
stoiskips spaces at the start and stops at the first character that cannot belong to the number. Only a field with no digits at the front, like""or"abc", makes it throw, as in mistake 3.Why build the report in an
ostringstreaminstead of printing as I go?Because some of the output depends on things you learn later, like David's line count. A string can also be printed twice, written to a file, or measured with
size(). For output you never need to hold, printing directly is fine.Would Program 6 line up Bangla names?
No.
setwpads by bytes, and lesson 01 showed that a Bangla letter is 3 bytes in UTF-8. A Bangla name would get too little padding and its columns would drift left. Lesson 04 names the library that counts letters instead of bytes.
Key takeaways
while (cin >> word)reads word by word, andsize()measures each word; the string grows to fit any word.vector<int> count(26, 0)withtolower(u) - 'a'counts letters, after casting eachchartounsigned char.- Split with
findandsubstr(start, length)when you need positions, or withgetline(in, field, ',')when you only need the fields. - A fresh
istringstreamper line reads the values inside it; an empty line simply reads none. - An
ostringstreamwithsetwandsetprecisionbuilds a whole report first, thenstr()prints it once. - Go deeper: Under the Hood, the small string buffer and the cost of + (Pro).
Next, lesson 04 asks when a string is the wrong choice, and draws the chart and the flowchart that pick a char, a string_view or a vector<char> instead.
End of lesson 3
Mark it done, and your progress moves with you.
Next: When to Use string and When Not To