Explain Regular Expression (Regex) and Its Types
A Regular Expression (Regex or RegExp) is a sequence of characters that forms a search pattern. It is used in Linux to find, match, and manipulate specific patterns in text. Regular expressions are commonly used with commands such as grep, sed, and awk, as well as editors like vi/vim.
Uses of Regular Expressions
Types of Regular Expressions
Linux supports two main types: Basic Regular Expression (BRE) and Extended Regular Expression (ERE).
BRE is the default regular expression syntax used by many Linux utilities such as grep and sed.
| Symbol | Meaning | Example |
|---|---|---|
| . | Matches any single character | b.t → bat, bit, but |
| ^ | Matches the start of a line | ^Hello |
| $ | Matches the end of a line | world$ |
| * | Matches zero or more occurrences of the previous character | lo* |
| [ ] | Matches one character from the given set | [aeiou] |
| [^ ] | Matches a character not in the given set | [^0-9] |
| \ | Escapes special characters | \. → literal . |
In BRE, characters such as +, ?, {}, and | generally need to be escaped to obtain their special meaning.
grep "^A" file.txt
This searches for lines beginning with A.
ERE provides additional and more convenient pattern-matching features. It is commonly used with grep -E, egrep, and sed -E.
| Symbol | Meaning | Example |
|---|---|---|
| + | One or more occurrences | lo+ |
| ? | Zero or one occurrence | colou?r |
| {n} | Exactly n occurrences | a{3} |
| {n,} | n or more occurrences | a{2,} |
| {n,m} | Between n and m occurrences | a{2,4} |
| | | Alternation (OR) | cat|dog |
| ( ) | Groups expressions | gr(a|e)y |
grep -E "cat|dog" animals.txt
This matches lines containing either cat or dog.
Difference Between BRE and ERE
| Feature | BRE | ERE |
|---|---|---|
| Type | Basic | Extended |
| +, ?, {} | Usually need escaping | Used directly |
| Grouping | \(\) | () |
| Alternation | \| | | |
| Common command | grep | grep -E, egrep |
Explain grep Command in Detail (Including fgrep, egrep)
grep stands for Global Regular Expression Print. It is a Linux command used to search for specific text patterns in files or input using regular expressions. Linux provides three commonly discussed variants:
grep is the basic pattern-matching command. It supports Basic Regular Expressions (BRE).
Syntax:
grep [options] pattern [file...]
Example:
grep "apple" file.txt
This displays the lines containing apple. Another example:
grep "^A" file.txt
This displays lines starting with A.
Important Features: supports BRE; matching is case-sensitive by default; can be combined with pipes and redirection; useful for searching patterns in one or more files.
Common grep Options
| Option | Purpose |
|---|---|
| -i | Ignores case while matching |
| -v | Displays lines that do not match |
| -n | Displays line numbers |
| -c | Counts matching lines |
| -l | Displays only filenames containing matches |
| -r / -R | Searches recursively in directories |
grep -i "linux" file.txt # Searches for linux, Linux, LINUX, etc.
grep -v "error" log.txt # Displays lines that do not contain error
egrep stands for Extended Global Regular Expression Print. It is equivalent to grep -E. It supports Extended Regular Expressions (ERE), allowing special characters such as +, ?, |, (), and {} to be used without escaping.
Syntax:
egrep [options] pattern [file...]
or
grep -E [options] pattern [file...]
Example:
egrep "cat|dog" animals.txt
This searches for lines containing either cat or dog. Another example:
egrep "a{2,4}" file.txt
This matches aa, aaa, or aaaa.
Main Advantage: egrep is useful for complex pattern matching because ERE provides additional metacharacters without requiring escape characters.
fgrep stands for Fixed String Global Regular Expression Print. It is equivalent to grep -F. Unlike grep and egrep, fgrep does not interpret regular-expression metacharacters. All characters are treated as literal text.
Syntax:
fgrep [options] string [file...]
or
grep -F [options] string [file...]
Example:
fgrep "a+b*c" file.txt
Here, + and * are treated as ordinary characters rather than regular-expression operators.
Main Advantage: fgrep is useful for exact or fixed-string searches and is generally faster for plain-text searching because it does not perform regular-expression parsing.
Difference Between grep, egrep and fgrep
| Feature | grep | egrep | fgrep |
|---|---|---|---|
| Full form | Global Regular Expression Print | Extended grep | Fixed grep |
| Regex support | Basic (BRE) | Extended (ERE) | No regex |
| Special characters | Some need escaping | Used directly | Treated as literals |
| Speed | Moderate | Slightly faster for complex patterns | Fastest for plain-text search |
| Best suited for | Simple patterns | Complex patterns | Exact string matching |
| Equivalent command | grep | grep -E | grep -F |
Explain sed Command in Detail
sed stands for Stream Editor. It is a powerful non-interactive text-processing utility used in Linux to process and transform text. It reads input line by line, applies the specified commands, and produces the processed output.
Unlike a normal text editor, sed can perform these operations directly on a text stream or file without manually opening and editing the file.
Basic Syntax
sed [options] 'command' filename
Example:
sed 's/old/new/' file.txt
This replaces the first occurrence of old with new on each line.
The s command is used for substitution.
Syntax:
sed 's/search_text/replacement_text/' file.txt
Example:
sed 's/apple/orange/' fruits.txt
This replaces the first occurrence of apple with orange on each line.
| Modifier | Meaning |
|---|---|
| s | Substitute |
| g | Replace all occurrences in each line |
| i | Perform case-insensitive search |
sed 's/apple/orange/g' fruits.txt # Replaces all occurrences of apple in each line
sed 's/APPLE/orange/i' fruits.txt # Performs a case-insensitive replacement
The d command is used to delete lines.
Syntax:
sed 'Nd' file.txt # N represents the line number
sed '2d' file.txt # Deletes the second line
sed '5,10d' file.txt # Deletes lines 5 through 10
sed '/error/d' file.txt # Deletes all lines containing the word error
The i command is used to insert text before a specified line.
Syntax:
sed 'N i\text' file.txt
Example:
sed '3 i\Inserted Line' file.txt
This inserts Inserted Line before line 3. It can also be used with a pattern:
sed '/pattern/ i\Before Match' file.txt
This inserts the specified text before a line matching the pattern.
The a command is used to append text after a specified line.
Syntax:
sed 'N a\text' file.txt
Example:
sed '2 a\Appended Line' file.txt
This adds Appended Line after line 2. It can also append text after a line matching a pattern:
sed '/pattern/ a\After Match' file.txt
Multiple sed commands can be executed using the -e option.
sed -e '1d' -e 's/old/new/' file.txt
This performs two operations: deletes the first line, and replaces old with new.
Useful sed Options
| Option | Purpose |
|---|---|
| -n | Suppresses automatic printing of lines |
| -e | Allows multiple sed commands |
| -i | Edits the file in-place, saving the changes directly |
sed -i 's/linux/ubuntu/g' file.txt
This replaces linux with ubuntu and saves the changes directly in the file.
Common Uses of sed
| Task | Command |
|---|---|
| Replace text | sed 's/linux/LINUX/g' file.txt |
| Delete first line | sed '1d' file.txt |
| Insert before line 2 | sed '2 i\Hello' file.txt |
| Append after matching lines | sed '/task/ a\Done' file.txt |