Sunday, July 11, 2021

Java Tutorial: Regular Expressions

Chapters

Java Regular Expressions

Regular Expressions(or regex) are sets of characters that form as patterns. These patterns can be used as criteria for searching specific sets of characters like words, names, repeating characters, etc.

There are two main classes that we need to understand in order to perform a search operation using regular expressions. The classes are: Matcher and Pattern class. These classes reside in java.util.regex package.

First off, we're going to compile our pattern by using compile() method then, we need to set our pattern and the input sequence(a character or a set of characters) for matching operation; We will use the matcher() method to do that. Then, We're going to choose a matching operation. There are three matching operations available for us: matches(), lookingAt() and find().

matches() Method

This method attempts to match the pattern to the entire input sequence. Input sequence is the string where the search operation is performed to. This method is one of the class members of Matcher class.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
  
    //we use Pattern.compile() to create a pattern
    //The string that is used as argument is the pattern
    //that we're going to use
    //Pattern.compile returns a Pattern object
    //with the created pattern
    Pattern pattern = Pattern.compile("Ghosts");
    
    //matcher() method sets the pattern and
    //input sequence to be ready for matching operation.
    //returns a Matcher object.
    //
    //
    //Once they are set, we can perform three matching
    //operations: matches(), lookingAt() and find()
    Matcher matcher = pattern.matcher("Brainy Ghosts Blogs");
    
    //
    //this will return false 'cause the pattern is just
    //a set of characters so, matches() method will just
    //attempt to match the entire pattern to the entire
    //region of the input sequence.
    System.out.println("Match Found: " + matcher.matches());
    
    //We can use method chaining to shortern our regex
    //syntax.
    boolean isMatch = Pattern.compile("Ghosts").
                      matcher("Ghosts").matches();
    
    //This will return true 'cause the pattern and the 
    //entire region of the input sequence are equal.
    System.out.println("Match Found: " + isMatch);
    
    //We can use the overloaded form of matches() method with
    //two parameters: the first parameter is the 
    //pattern and the second parameter is the input sequence.
    isMatch = Pattern.matches("Ghosts","Ghosts Blogs");
    
    //this will return false
    System.out.println("Match Found: " + isMatch);
    
    //The compile() method has overloaded method with two
    //paramaters. The first parameter is the pattern and 
    //the second is a flag. There are different types of
    //flags that we can use as argument. To know the flags,
    //visit the Pattern class documentation.
    //
    //Pattern.CASE_INSENSITIVE ignores the lettercases of both
    //pattern and the input sequence during comparison.
    //
    //Two add multiple flags, use the bitwise OR operator("|")
    //e.g. Pattern.compile("Ghosts",Pattern.CASE_INSENSITIVE | 
    //                              Pattern.MULTILINE);
    //
    //or put the flags in an int variable
    //e.g
    //int flags = Pattern.CASE_INSENSITIVE | Pattern.MULTILINE;
    //Pattern.compile("Ghosts",flags);
    pattern = Pattern.compile("Ghosts",Pattern.CASE_INSENSITIVE);
    matcher = pattern.matcher("GhOsTs");
    
    //This will return true regardless of the lettercases of
    //the pattern and input sequence.
    System.out.println("Match Found: " + matcher.matches());
    
  }
}
lookingAt() Method

This method attempts to match the pattern against the region of the input sequence, Starting from the first index(0). This is similar to matches(), but unlike matches(), this method is not required to match the entire region of the input sequence against the pattern. This method is one of the class members of Matcher class.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    Pattern pattern = Pattern.compile("Ghosts");
    Matcher matcher = pattern.matcher("Ghosts "
                                     +"Blogs Ghosts");
                                 
    if(matcher.lookingAt()){
      System.out.println("Match Found!");
      System.out.println("Start index: " + matcher.start());
      System.out.println("End index: " + matcher.end());
      System.out.println();
    }
  }
}
find() Method

This method scans the input sequence and look for each subsequence that matches the pattern. We will use the start() and end() methods to determine the starting and end indexes of the matched subsequences(subsequence is a sequence that is a subset of a sequence) starting from the first index which is 0. This method is one of the class members of Matcher class.
Note: start() returns the start index of the previous match whereas end() returns the offset after the last character matched.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    Pattern pattern = Pattern.compile("Ghosts");
    Matcher matcher = pattern.matcher("Brainy Ghosts "
                                     +"Blogs Ghosts");
    int matchCount = 0;
    
    //find() method will look for each subsequence
    //that matches the pattern even a match is already
    //found as long as the matcher doesn't reset.
    //Otherwise, find() method starts at the
    //beginning of the input sequence
    while(matcher.find()){
      matchCount++;
      System.out.println("Match Found!");
      System.out.println("Start index: " + matcher.start());
      System.out.println("End index: " + matcher.end());
      System.out.println();
    }
    System.out.println("Match Count: " + matchCount);
    System.out.println();
    System.out.println("Resetting Matcher...");
    //reset the matcher
    matcher.reset();
    
    //we can use find() method to get the first match
    //only.
    if(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("Start index: " + matcher.start());
      System.out.println("End index: " + matcher.end());
      System.out.println();
    }
    
  }
}
split() Method

This method splits the input sequence into multiple sequences, which is based on a delimiter. This method is one of the class members of Pattern class.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Apple,Banana,Citrus,Durian";
    
    //"," is the delimiter that we're going to use
    Pattern pattern = Pattern.compile(",");
    String[] strArr = pattern.split(str);
    
    for(String s: strArr)
      System.out.println(s);
  }
}
replaceAll() Method

This method replaces the subsequence(sequence that is a subset of a sequence) that matches the pattern with the replacement string and returns the constructed string with the string that replaces the subsequence. This method is one of the class members of Matcher class.
Method form: replaceAll(String replacement){}

import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Apple,Banana,Citrus,Durian";
  
    Pattern pattern = Pattern.compile(",");
    Matcher matcher = pattern.matcher(str);
    
    String modifiedStr = matcher.replaceAll("-");
    System.out.println(modifiedStr);
    System.out.println(str);
  }
}
region() method

This method limits the searchable region of an input sequence. By default, matching operations use the entire region of an input sequence if necessary. By using region() method, we can specify which region is searchable.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "apt, ape, apart, uptown apostle";
    Pattern pattern = Pattern.compile("ap");
    Matcher matcher = pattern.matcher(str);
    
    //the first parameter is the starting point
    //of our searchable region
    //the second parameter is the end point of
    //our searchable region
    matcher.region(0,9);
    
    while(matcher.find()){
      System.out.println("Found a Match!");
      System.out.println("First index: " + matcher.start());
      System.out.println("Last index: " + matcher.end());
      System.out.println();
    }
  }
}

Result

Found a Match!
First index: 0
Last index: 2

Found a Match!
First index: 5
Last index: 7
PatternSyntaxException

PatternSyntaxException is an exception being thrown if we use an invalid pattern.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    try{
      //The pattern here is invalid. Therefore,
      //java will throw a PatternSyntaxException
      //at runtime.
      Pattern pattern = Pattern.compile("\\");
    }
    catch(PatternSyntaxException e){
      e.printStackTrace();
    }
    
  }
}
Regular-Expression Constructs

We already know that we can use string literals as a construct in a pattern. Regular expression is not limited to string literals, There are other constructs that we can use to create a pattern.

Characters

We can use characters in octal,ascii and unicode format; Also, we can use special escape sequence like "\n" as part of our pattern. We need to wrap characters into double quotes when using as a pattern or part of a pattern.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "\"Banana\" \"Nuts\"";
  
    //\42 is the octal representation of double
    //quotes(")
    Pattern pattern = Pattern.compile("\042");
    Matcher matcher = pattern.matcher(str);
    int quotesCount = 0;
    
    while(matcher.find())
      quotesCount++;
    
    System.out.println("Input: " + str);
    System.out.println("Double quotes count: " + quotesCount);
    
    str = "Banana\nNuts\n";
    //\n or newline is a special escape sequence
    pattern = Pattern.compile("\n");
    String[] strArr = pattern.split(str);
    
    for(String s : strArr)
      System.out.println(s);
  }
}
Character Classes

Character classes are sets of characters enclosed within brackets([]). Character classes may contain union(implicit) or intersection(&&) operator. Character classes may also contain another character class and other operators like the negate(^) and range(-) operators.
Note: Java standard operators and regex operators function differently.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    
    //Using lookingAt() operation
    //[tuv] means, one of these letters(t,u,v)
    //must be the first character in the
    //input sequence for a match to happen.
    Pattern pattern = Pattern.compile("[tuv]",
                      Pattern.CASE_INSENSITIVE);
    Matcher matcher = pattern.matcher("Tell Me Why...");
    boolean isMatch = matcher.lookingAt();
    System.out.println("using lookingAt()");
    System.out.println("Match Found? " + isMatch);
    System.out.println();
    
    //reset matcher with new input sequence
     matcher.reset("abc");
    
    //Using matches() operation
    //[tuv] means,only one of these letters(t,u,v)
    //must be in the input sequence for a
    //match to happen.
    //It's not considered as a match if two
    //or more of those letters(t,u,v) are
    //present in the input sequence
    //
    //matches() will return false 'cause
    //the input sequence is not "t","u"
    //or "v"
    isMatch = matcher.matches();
    System.out.println("using matches()");
    System.out.println("Match Found? " + isMatch);
    System.out.println();
    
    //reset matcher with new input sequence
    matcher.reset("t");
    //This will return true 'cause the input
    //sequence is one of the letters of the
    //character class([tuv])
    isMatch = matcher.matches();
    System.out.println("using matches()");
    System.out.println("Match Found? " + isMatch);
    System.out.println();
     
    //reset matcher with new input sequence
    matcher.reset("tu");
    
    //This will return false 'cause the input
    //sequence is not "t","u" or "v".
    isMatch = matcher.matches();
    System.out.println("using matches()");
    System.out.println("Match Found? " + isMatch);
    System.out.println();
    
    //Using find() operation
    //[tuv] means, one of these letters(t,u,v)
    //must be in the input sequence for a
    //match to happen.
    
    String str = "Tell them to vacate the area and"+
                 " hide underground!";
    //reset matcher with new input sequence
    matcher.reset(str);
    
    System.out.println("using find()");
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("index: " + matcher.start());
      System.out.println("Character: " + 
                         str.charAt(matcher.start()));
      System.out.println();
    }
    
  }
}
Negate(^)

Let's create another example. Let's use the negate(^) operator this time. Negate(^) operator reversed the effect of an expression in a character class.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Trace my card.";
  
    //[^td] means, any character except for "t"
    //and "d".
    Pattern pattern = Pattern.compile("[^td]",
                      Pattern.CASE_INSENSITIVE);
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Input Sequence: " + str);
    System.out.print("Result: ");
    while(matcher.find())
      System.out.print(str.charAt(matcher.start()));
  }
}
Range(-)

Next, let's use the range(-) operator. Range(-) operator simply sets a range between two characters.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Gone, Igloo, Fear, Hall";
    
    //[f-i] means, any letter between "f" and "i"
    //in this example, we use the negate operator
    //so, [^f-i] means, any letter except for the
    //letters between "f" and "i"
    Pattern pattern = Pattern.compile("[^f-i]",
                      Pattern.CASE_INSENSITIVE);
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Input Sequence: " + str);
    System.out.print("Result: ");
    while(matcher.find())
      System.out.print(str.charAt(matcher.start()));
  }
}
Note, the left operand of range operator should be greater than the right operand. Otherwise, you will encounter a PatternSyntaxException. Try changing [^f-i] to [^i-f].

Here's a list of unicode characters. Each character has a number representation, you can use those numbers as reference to compare if a character is greater than the other one.

Union and Intersection(&&)

Next, let's try the union(implicit) and intersection(&&) operators. Quoted from Pattern class documentation: "The union operator denotes a class that contains every character that is in at least one of its operand classes. The intersection operator denotes a class that contains every character that is in both of its operand classes."

First off, Let's crete an example to demonstrate union operator.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Six6Nine9";
    
    //This pattern "[a-z[A-Z]]" has a union operator
    //between "a-z" and "[A-Z]" expressions.
    //
    //[a-z[A-Z]] can be shortened into [a-zA-Z]
    //
    //[a-z[A-Z]] means, any character ranging from "a"
    //to "z" or "A" to "Z"
    Pattern pattern = Pattern.compile("[a-z[A-Z]]");
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Input Sequence: " + str);
    System.out.print("Result: ");
    while(matcher.find())
      System.out.print(str.charAt(matcher.start()));
  }
}
Next, let's create an example to demonstrate intersection operator.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "because dude";
    
    //"[a-d&&[bc]]" can be shortened into:
    //"[a-d&&bc]"
    Pattern pattern = Pattern.compile("[a-d&&[bc]]");
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Input Sequence: " + str);
    System.out.print("Result: ");
    while(matcher.find())
      System.out.print(str.charAt(matcher.start()));
  }
}
The result is "bc" because the characters "b" and "c" are present in both intersection operands(a-d and bc).

Let's try a more complex character class.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "1234abcdefghi";
    
    //This pattern means, "123" or "a" to "z" but limited to
    //"a" to "g" and exclude "bc".
    //The result is: 123adefg
    //4 is not included in the result 'cause we only include
    //"123" in our pattern.
    //b and c are not included 'cause those are excluded
    //h and i are not included 'cause they're not in the 
    //range between "a" and "g"
    String regEx = "[[123][a-z&&[a-g&&[^bc]]]]";
    Pattern pattern = Pattern.compile(regEx);
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Input Sequence: " + str);
    System.out.print("Result: ");
    while(matcher.find())
      System.out.print(str.charAt(matcher.start()));
  }
}
Pre-Defined Character Classes

The are characters or metacharacters that denote sets of pre-defined character classes. For example, "\d" denotes a digit-only character class "[0-9]", "." denotes any character(may or may not match line terminators), etc. Check out Pattern class documentation: for more pre-defined character classes.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    
    System.out.println("\".\" metacharacter");
    //"p.." means, a "p" and two characters of
    //any kind
    boolean isMatch = Pattern.matches("p..","pie");
    //This will return true.
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches("p..","pit");
    //This will return true.
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches("p..","pi");
    //This will return false 'cause the match
    //operation requires three characters.
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches("p..","sit");
    //This will return false 'cause the first
    //character must be "p"
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches(".pe","ape");
    //This will return true 'cause the first
    //character can be any character
    System.out.println("isMatch? " + isMatch);
    
    //"." in the character class brackets and
    //"." outside the brackets function differently
    //"." outside brackets works as a metacharacter
    //whereas "." inside brackets works as
    //a literal.
    //So, .[.] means, any character followed by
    //a dot.
    isMatch = Pattern.matches(".[.]","a.");
    //This will return true
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches(".[.]","ap");
    //This will return false
    System.out.println("isMatch? " + isMatch);
    System.out.println();
    
    System.out.println("\\d metacharacter");
    //\d is equivalent to [0-9] character class
    //We need to escape "\" by using "\"
    //that's why there are two "\"
    isMatch = Pattern.matches("\\d","2");
    //This will return true
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches("\\d","a");
    //This will return false
    System.out.println("isMatch? " + isMatch);
    System.out.println();
    
    System.out.println("\\W metacharacter");
    //\W means whitespace character. \W is
    //equivalent to [ \t\n\x0B\f\r] character class
    isMatch = Pattern.matches("\\W"," ");
    //This will return true
    System.out.println("isMatch? " + isMatch);
    
    isMatch = Pattern.matches("\\W","\t");
    //This will return true
    System.out.println("isMatch? " + isMatch);
  }
}
Note: Some metacharacters that denote pre-defined character classes may not work as intended if they're inside the chracter class brackets([]). One example is the "." metacharacter that I explain in the example above.

Boundary Matchers

boundary matchers are metacharacters that can be used to create a pattern regarding bounds. The metacharacters are placed at the beginning or end of a pattern. For example, "^" character denotes that a character or set of characters must be in the beginning of a line for a match to happen e.g. ^Hello. Check out Pattern class documentation: for more boundary matchers.

^ and $

"^" matches the input sequence at the beginning of the line whereas "$" matches the input sequence at the end of the line.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Hello This is Me!\nHello There!";
    
    Pattern pattern = Pattern.compile("^Hello");
    Matcher matcher = pattern.matcher(str);
    int helloCount = 0;
    
    while(matcher.find()) helloCount++;
    
    System.out.println("Input: " + str);
    //helloCount is 1 'cause there's one "Hello" word
    //at the beginning
    System.out.println("How many Hello: " + helloCount);
    System.out.println();
    
    matcher.reset();
    helloCount = 0;
    
    //"$" denotes that a character or a word
    //must be at the end of a line.
    matcher.usePattern(Pattern.compile("Hello$"));
    
    while(matcher.find()) helloCount++;
    
    System.out.println("Input: ");
    System.out.println(str);
    //helloCount is 0 'cause there's no "Hello" word
    //at the end
    System.out.println("How many Hello: " + helloCount);
    System.out.println();
    
    boolean isMatch = Pattern.matches("^Hello$","Hello");
    //This will return true 'cause the word "Hello" is
    //at the beginning and end of the input sequence
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("^Hello$",str);
    //This will return false 'cause "Hello" is not
    //at the beginning and at the end of the input
    //sequence
    System.out.println("isMatch: " + isMatch);
  }
}
In the example above, java considers "Hello This is Me!\nHello There!" input as a single line even there's a \n in the input because by default, java ignores line terminators(\n, \r\n, etc.) if we use "^" or "$" metecharacters. To allow java to recognize line terminators when using "^" or "$", we need to enable the MULTILINE flag.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Hello This is Me!\nHello There!";
    
    //Enable MULTILINE Flag
    Pattern pattern = Pattern.compile("^Hello",Pattern.MULTILINE);
    Matcher matcher = pattern.matcher(str);
    int helloCount = 0;
    
    while(matcher.find()) helloCount++;
    
    System.out.println("Input: " + str);
    //helloCount is 2 'cause java recognizes \n so,
    //"Hello This is Me!" is on the first line and
    //"Hello There!" is on the second line.
    System.out.println("How many Hello: " + helloCount);
    System.out.println();
  }
}
Word Boundary

Next, let's use the word boundary(\b).Word boundary matches the first character or the character after the last character of a word in the input sequence.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
  
    String str = "Hello! Java,regex File?Tutorial";
	
	Pattern pattern = Pattern.compile("\\b");
    Matcher matcher = pattern.matcher(str);
    
    while(matcher.find()){
      try{
        System.out.println("index: " + matcher.start() +
                           " character: " + 
                           str.charAt(matcher.start()));
      }
      catch(StringIndexOutOfBoundsException e){
        System.out.println("An exception occured");
        System.out.println("Input last index: " + 
                          (str.length()-1));
        System.out.println("Matcher index: " + matcher.start());
      }
    }
    
  }
}

Result

index: 0 character: H
index: 5 character: !
index: 7 character: J
index: 11 character: ,
index: 12 character: r
index: 17 character: 
index: 18 character: F
index: 22 character: ?
index: 23 character: T
An exception occured
Input last index: 30
Matcher index: 31
So, the first match is "H" which is the first character of "Hello" word then, the next match is "!" which is the character after "o" which is the last character of "Hello" word and so on. Notice the last match, Matcher index is larger than the last index of the input sequence. Index 31 is an ambiguous match and it's called "zero-length match".

Try removing the try-catch clause in the example above and you will encounter an StringOutOfBoundsException in the last part of the matching operation.

We can specify which word boundary to get by adding a sequence next to "\b" or after it.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
  
    String str = "Hello! Java,regex Jill?Loo";
	
	Pattern pattern = Pattern.compile("\\bJ");
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Get first subsequence boundary");
    while(matcher.find()){
      if(matcher.start() < str.length())
        System.out.println("index: " + matcher.start() +
                           " character: " + 
                           str.charAt(matcher.start()));
      else
        System.out.println("An empty string");
    }
    
    System.out.println();
    matcher.reset();
    matcher.usePattern(Pattern.compile("o\\b"));
    
    System.out.println("Get last subsequence boundary");
    while(matcher.find()){
      if(matcher.start() < str.length())
        System.out.println("index: " + matcher.start() +
                           " character: " + 
                           str.charAt(matcher.start()));
      else
        System.out.println("An empty string");
    }
    
  }
}

Result

Get first subsequence boundary
index: 7 character: J
index: 18 character: J

Get last subsequence boundary
index: 4 character: o
index: 25 character: o
Non-Word Boundary

Non-word boundary is like the opposite of word boundary. Non-word boundary is a boundary between two characters(word or non-word). I like to think of word boundary as container bounds whereas non-word boundary as separator bounds, If you will.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Hello! This is me!";
	
	Pattern pattern = Pattern.compile("\\B");
    Matcher matcher = pattern.matcher(str);
    
    String result = "";
    while(matcher.find())
      try{
        result += str.charAt(matcher.start());
      }catch(StringIndexOutOfBoundsException e){
        System.out.println("An exception occured");
        System.out.println("Input last index: " + 
        (str.length()-1));
        System.out.println("Matcher index: " + matcher.start());
      }
      System.out.println();
      System.out.println("Non-word boundary");
      System.out.println("Input: " + str);
      System.out.println("result: " + result);
      System.out.println();
      
      pattern = Pattern.compile("\\b");
      matcher = pattern.matcher(str);
      
      result = "";
      while(matcher.find())
      try{
        result += str.charAt(matcher.start());
      }catch(StringIndexOutOfBoundsException e){
        System.out.println("An exception occured");
        System.out.println("Input last index: " + (str.length()-1));
        System.out.println("Matcher index: " + matcher.start());
      }
      System.out.println();
      System.out.println("Word boundary");
      System.out.println("Input: " + str);
      System.out.println("result: " + result);
      System.out.println();
      
  }
}

Result

An exception occured
Input last index: 17
Matcher index: 18

Non-word boundary
Input: Hello! This is me!
result: ello hisse


Word boundary
Input: Hello! This is me!
result: H!T i m!
So, In the non-word boundary result, "ello" sequence is the bounds(separator) between "H" and "!" characters, whereas in the word boundary result, "H" and "!" are the bounds(container) of "ello" sequence. "his" sequence is the bounds(separator) between "T" and " "(whitespace) characters whereas "T" and " " are bounds(container) of "his" sequence and so on.

We can specify which non-word boundary to get by adding a sequence next to "\B" or after it.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String str = "Hello! meat is great!";
	
	Pattern pattern = Pattern.compile("\\Be");
    Matcher matcher = pattern.matcher(str);
    
    System.out.println("Pattern: " + "\\Be");
    System.out.println("Input: " + str);
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("index: " + matcher.start());
    }
    System.out.println();
    
    pattern = Pattern.compile("a\\B");
    matcher = pattern.matcher(str);
    
    System.out.println("Pattern: " + "a\\B");
    System.out.println("Input: " + str);
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("index: " + matcher.start());
    }
  }
}

Result

Pattern: \Be
Input: Hello! meat is great!
Match Found!
index: 1
Match Found!
index: 8
Match Found!
index: 17

Pattern: a\B
Input: Hello! meat is great!
Match Found!
index: 9
Match Found!
index: 18
Quantifiers

Note: Quantifiers can be used in conjunction with character classes and capturing groups.
As the name implies, Quantifiers quantifies the quantity of how many matches a sequence has in the input sequence. There are three types of quantifiers: Greedy, Reluctant and Possessive. These quantifiers have similar constructs, However, Their mechanics are subtly different. Let's create an example to demonstrate quantifiers.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
  	System.out.println("? quantifier");
    //"?" means once or not at all.
    boolean isMatch = Pattern.matches("pie?","pie");
    //isMatch is true 'cause "e" occurs only once
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("pie?","pi");
    //isMatch is true 'cause "e" doesn't occur
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("pie?","piee");
    //isMatch is false 'cause "e" occurs many times
    System.out.println("isMatch: " + isMatch);
    System.out.println();
    
    System.out.println("* quantifier");
    //"*" means zero or more
    isMatch = Pattern.matches("pi*e","pie");
    //isMatch is true 'cause "i" occurs one time
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("pi*e","pe");
    //isMatch is true 'cause "i" doesn't occur
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("pi*e","piiie");
    //isMatch is true 'cause "i" occurs many times
    System.out.println("isMatch: " + isMatch);
    System.out.println();
    
    System.out.println("+ quantifier");
    //"+" means one or more
    isMatch = Pattern.matches("p+ie","pie");
    //isMatch is true 'cause "p" occurs one time
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("p+ie","pppie");
    //isMatch is true 'cause "p" occurs many times
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("p+ie","ie");
    //isMatch is false 'cause "p" doesn't occur
    System.out.println("isMatch: " + isMatch);
    System.out.println();
    
    System.out.println("{n} quantifier");
    //{n} means, exactly n times
    isMatch = Pattern.matches("p{2}ie","ppie");
    //isMatch is true 'cause "p" occurs
    //exactly two times
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("p{2}ie","pie");
    //isMatch is false 'cause "p" occurs one time
    System.out.println("isMatch: " + isMatch);
    System.out.println();
    
    System.out.println("{n,} quantifier");
    //{n,} means, at least n times
    isMatch = Pattern.matches("p{2,}ie","pppie");
    //isMatch is true 'cause "p" occurs
    //at least two times
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("p{2,}ie","pie");
    //isMatch is false 'cause "p" doesn't occur
    //at least two times
    System.out.println("isMatch: " + isMatch);
    System.out.println();
    
    System.out.println("{n,m} quantifier");
    //{n,m} means, at least n times and not more
    //than m times
    isMatch = Pattern.matches("pie{2,3}","piee");
    //isMatch is true 'cause "e" occurs at
    //least 2 times and not more than 3 times 
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("pie{2,3}","pieeee");
    //isMatch is false 'cause "e" occurs 
    //more than 3 times 
    System.out.println("isMatch: " + isMatch);
    
    isMatch = Pattern.matches("pie{2,3}","pie");
    //isMatch is false 'cause "e" doesn't occur 
    //at least 2 times 
    System.out.println("isMatch: " + isMatch);
    System.out.println();
    
    //This pattern is more complicated than previous
    //patterns and we need to understand the subtle
    //differences between greedy, reluctant and
    //possessive quantifiers in order to deeply
    //understand this pattern
    Matcher matcher = Pattern.compile(".*pie").
                      matcher("apple pie apple pie");
    //Result:
    //Match Found!
    //start index: 0
    //end index: 19
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("start index: " + matcher.start());
      System.out.println("end index: " + matcher.end());
    }
    
  }
}
Greedy Quantifiers

Greedy quantifiers grabs every character it can grab in the input for itself and then try to match the input against the subsequent part. The matcher will release one character from the grabbed sequence back to the input as the operation progresses, starting from the last index. The matcher will stop releasing characters if the overall pattern is satisfied or there's no character to be released.
Greedy Quantifiers
X? 	X, once or not at all
X* 	X, zero or more times
X+ 	X, one or more times
X{n} 	X, exactly n times
X{n,} 	X, at least n times
X{n,m} 	X, at least n but not more than m times
Let's create an example to demonstrate the mechanics of greedy quantifier.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
	
    //".*" means one or more character/s of
    //any kind before the sequence "pie"
    //"." means any character, "*" is a greedy
    //quantifier means zero or more occurences
	Pattern pattern = Pattern.compile(".*pie");
    Matcher matcher = pattern.matcher("apple pie apple pie");
    
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("start index: " + matcher.start());
      System.out.println("end index: " + matcher.end());
    }
  }
}

Result

Match Found!
start index: 0
end index: 19
So first off, we have two parts here: ".*" and "pie". ".*" grabs the entire sequence in the input first and then the matcher tries to match the input against the subsequent part. The problem here is that ".*" grabs the entire sequence in the input so, we don't have any sequence to be matched against.

The overall pattern needs to be satisfied for a match to happen. So, the matcher releases one character from the grabbed sequence, starting from the last index and then tries to match the relased characters against the subsequent part.

If the match is still a failure then, the matcher will release another character until the overall pattern is satisfied or there's no character to be released. ".*pie" will be satisfied if the released characters are equivalent to "pie" and ".*" holds a character or nothing.

The start index is 0 'cause that's the starting index of the most suitable match. End index is 19 'cause end() method returns the offset after the last character matched. The lotal length of the input sequence is 18 and that last character in index 19 is an ambiguous match which is called "zero-length match".

Reluctant Quantifiers

Reluctant quantifier doesn't grab every character it can grab in the input for itself immediately, it lets the sequence in the input to be matched against the subsequent part first. Then, grabs one character as the matching operation progresses, starting from the first index. The quantifier will stop grabbing characters if the overall pattern is satisfied or there's no character to be grabbed.
Reluctant quantifiers
X?? 	X, once or not at all
X*? 	X, zero or more times
X+? 	X, one or more times
X{n}? 	X, exactly n times
X{n,}? 	X, at least n times
X{n,m}? 	X, at least n but not more than m times
Let's create an example to demonstrate the mechanics of reluctant quantifier.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
	
    //Reluctant quantifier construct is similar to
    //greedy but with additional "?" before the
    //quantifier.
    //
    //".*?" still accepts zero or more
    //characters of any kind, but this time, the
    //quantifier is reluctant.
	Pattern pattern = Pattern.compile(".*?pie");
    Matcher matcher = pattern.matcher("apple pie apple pie");
    
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("start index: " + matcher.start());
      System.out.println("end index: " + matcher.end());
      System.out.println();
    }
  }
}

Result

Match Found!
start index: 0
end index: 9

Match Found!
start index: 9
end index: 19
First off, the input is matched against "pie" which fails 'cause the first sequence of the input is not equal to "pie". So, the quantifier grabs one character from the starting index and then, match the input against the subsequent part until it finds an overall match. When our search reaches index 6 the overall pattern is satisfied. Now, start() and end() give the first and last index of the previous match.

Since we're using find(), our search operation continues until it reaches the second match with a start index of 9 and end index of 19.

Possessive Quantifiers

When we use "." with a possessive quantifier then, it grabs every character it can grab in the input. However, possessive quantifier doesn't release any characters that it grabbed.
Possessive quantifiers
X?+ 	X, once or not at all
X*+ 	X, zero or more times
X++ 	X, one or more times
X{n}+ 	X, exactly n times
X{n,}+ 	X, at least n times
X{n,m}+ 	X, at least n but not more than m times
Let's create an example to demonstrate the mechanics of possessive quantifier.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
	
    //Possessive quantifier construct is similar to
    //greedy but with additional "+" before the
    //quantifier.
    //
    //".*+" still accepts zero or more
    //characters of any kind, but this time, the
    //quantifier is possessive.
	Pattern pattern = Pattern.compile(".*+pie");
    Matcher matcher = pattern.matcher("apple pie apple pie");
    
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("start index: " + matcher.start());
      System.out.println("end index: " + matcher.end());
      System.out.println();
    }
    System.out.println("End of the program.");
    
  }
}

Result

End of the program.
As we can see, our program didn't find any matches that's because ".*+" didn't release any characters back to the input to be matched against "pie". So, the matcher can't find a match between input and the overall pattern.

Use possessive quantifier if you want to grab all input characters without needing to release any of those. It will outperform the equivalent greedy quantifier in a situation where a match is not immediately found.

Zero-Length Match

Zero-length match happens when a match that is found is ambiguous. We already encountered zero-length matches in previous topics. These ambiguous matches can be found in an empty input string(""); beginning and end of an input string; between two characters of an input string. Zero-length match can be spotted easily by checking the start and end index of a match. If both indexes are equal then, that match is a zero-length match.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    Pattern pattern = Pattern.compile("a?");
    Matcher matcher = pattern.matcher("");
    
    System.out.println("a? pattern");
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("start index: " + matcher.start());
      System.out.println("end index: " + matcher.end());
      System.out.println();
    }
    
    matcher = Pattern.compile("a*").matcher("");
    
    System.out.println("a* pattern");
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.println("start index: " + matcher.start());
      System.out.println("end index: " + matcher.end());
      System.out.println();
    }
    
  }
}

Result
Match Found!
start index: 0
end index: 0

Match Found!
start index: 0
end index: 0
The matcher found a match even we use an empty string. This happens 'cause "?" and "*" accepts zero occurences. "?" and "*" checks if "a" appears in the empty string, a match happens if "a" is found or not. Try this pattern "a+" and the matcher won't find any matches in the example above. Because, "+" doesn't accept zero occurences. Let's try another example.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    Pattern pattern = Pattern.compile("a?");
    Matcher matcher = pattern.matcher("abaaa");
    
    System.out.println("pattern: a?");
    System.out.println("input: abaaa");
    System.out.println("Matches");
    while(matcher.find()){
      System.out.print("start index: " + matcher.start());
      System.out.print(" ");
      System.out.print("end index: " + matcher.end());
      System.out.println();
    }
    
    matcher = Pattern.compile("a*").matcher("abaaa");
    
    System.out.println();
    System.out.println("pattern: a*");
    System.out.println("input: abaaa");
    System.out.println("Matches");
    while(matcher.find()){
      System.out.print("start index: " + matcher.start());
      System.out.print(" ");
      System.out.print("end index: " + matcher.end());
      System.out.println();
    }
    
    matcher = Pattern.compile("a+").matcher("abaaa");
    
    System.out.println();
    System.out.println("pattern: a*");
    System.out.println("input: abaaa");
    System.out.println("Matches");
    while(matcher.find()){
      System.out.print("start index: " + matcher.start());
      System.out.print(" ");
      System.out.print("end index: " + matcher.end());
      System.out.println();
    }
    
  }
}

Result
a? pattern
input: abaaa
Matches
start index: 0 end index: 1
start index: 1 end index: 1
start index: 2 end index: 3
start index: 3 end index: 4
start index: 4 end index: 5
start index: 5 end index: 5

a* pattern
input: abaaa
Matches
start index: 0 end index: 1
start index: 1 end index: 1
start index: 2 end index: 5
start index: 5 end index: 5

a+ pattern
input: abaaa
Matches
start index: 0 end index: 1
start index: 2 end index: 5
If we look at the result, "a?" checks every character 'cause "?" means once or not at all. So, "a?" needs to check if "a" appears or not in every character. "a*" checks the input differently, "*" means zero or more. So, First, "a*" checks if "a" is present, if it's true then, it checks if this "a" is followed by another a and so on.

Notice this result "start index: 1 end index: 1" in both "?" and "*" quantifiers. Index 1 has "b" character in it and it's considered as a match. This happens 'cause "?" and "*" only check if "a" appears in that index or not. "?" and "*" don't check the difference between "a" and "b". If "a" doesn't appear in those indexes then, the result is zero occurence or zero-length match.

Next, look at the result "start index: 5 end index: 5" in both "?" and "*" quantifiers. The max index of the "abaaa" is 4, but still, we got a character at index 5. Index 5 doesn't exist in the index range of the input, therefore the character in index 5 is an empty string. Thus, the result is a zero-length match. Though, it could be a null character that terminates a string.

Now, let's check the "a+" pattern. "+" means one or more. "+" checks if "a" appears once in an index then checks if it's followed by another "a" and so on. In the result, "+" didn't consider index 1 and index 5 as valid matches 'cause "a" didn't appear in those indexes at least once.

Logical Operators

We use logical operators to combine different types of regular-expression constructs. These are the logical operators:
XY(AND operator)
X|Y(OR operator)
(X)(Capturing Group)
We're going to discuss AND and OR operators in this topic. Capturing group has its own topic. Let's create an example to demonstrate AND and OR operators.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    
    //.+[mu].++ simpily means, one or more
    //characters on both sides of "m" or "u".
    //".+" is a greedy quantifier whereas .++
    //is a possessive quantifier.
    //
    //This is a combination of both quantifiers
    //and character class. AND operator is 
    //implicit, That's why we don't see any
    //operator token between ".+", "[mu]" and ".++"
    //
    //regex constructs that are connected with one
    //another using AND operator are combined
    //together and each construct must be satisfied
    //in the given order for a match to happen
    Matcher matcher = Pattern.compile(".+[mu].++").
                      matcher("immutable");
    
    boolean isMatch = matcher.matches(); 
    //isMatch is true
    System.out.println("AND operator");
    System.out.println("isMatch? " + isMatch);
    
    matcher.reset("metabolism");
    isMatch = matcher.matches();
    //isMatch is false
    System.out.println("isMatch? " + isMatch);
    System.out.println();
    
    //In this pattern, we use "|" or OR operator
    //between ^facebook$ and ^youtube$".
    //
    //regex constructs that are connected with one
    //another using OR operator are combined
    //together and only one construct is required
    //to be satisfied for a match to happen
    //
    //Don't be confused between "^facebook$|^youtube$"
    //and "^facebook|youtube$". These two are different.
    //
    //"^facebook|youtube$" means "facebook" at the start
    //of a line or "youtube" at the end of a line.
    //
    //"^facebook$|^youtube$" means, "facebook" or 
    //"youtube" at the start and end of a line.
    matcher = Pattern.compile("^facebook$|^youtube$").
              matcher("facebook");
    
    isMatch = matcher.matches();      
    System.out.println("OR operator");
    //isMatch is true 'cause even ^youtube$
    //is not satisfied, ^facebook$
    //is satisfied.
    //
    //Since we're using OR operator between
    //^facebook$ and ^youtube$, only one of
    //them is needed to be satisfied for a
    //match to happen
    System.out.println("isMatch? " + isMatch);
              
  }
}
Escaping Metacharacters

Metacharacters are characters that have special meaning in regular expressions like brackets([]), braces({}), dot(.), etc. Sometimes, we wanna override their intended purposes. Escaping metacharacters is similar to escaping character sequence, we will use backslash(\) to escape metacharacters.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    
    //use double backslash to escape "[" and "]"
    //if we only use one backslash then, the compiler
    //will consider "[" and "]" as special characters like \n
    //which is invalid 'cause "\[" and "\]" are 
    //not special characters.
    //
    //we add another backslash to escape the backslash
    //after the newly added backslash. Now, the matcher
    //treats those metacharacters as regular literals.
    //
    //In this pattern, "[" and "]" are treated as
    //literals
    boolean isMatch = Pattern.matches("\\[\\]","[]");
    //isMatch is true
    System.out.println("isMatch? " + isMatch);
    
    //This pattern escapes "." which means
    //"any characters"
    Matcher matcher = Pattern.compile("\\.+").
                      matcher(".abba.");
    
    System.out.println();
    //Two matches are found 'cause "\\.+" looks
    //for one or more occurences of dot(.)
    System.out.println("Matches");
    while(matcher.find()){
      System.out.print("start index: " + matcher.start());
      System.out.print(" ");
      System.out.print("end index: " + matcher.end());
      System.out.println();
    }
    
    //In this pattern, "." is not escaped
    //matcher treats the "." here as a
    //metacharacter
    matcher = Pattern.compile(".+").
                      matcher(".abba.");
    
    System.out.println();
    //match index is ranging from 0 to 6 
    //'cause ".+" looks for one or more
    //occurences of any character
    System.out.println("Matches");
    while(matcher.find()){
      System.out.print("start index: " + matcher.start());
      System.out.print(" ");
      System.out.print("end index: " + matcher.end());
      System.out.println();
    }
  }
}
Next, let's escape metacharacters with backslash like "\b" and others.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
  
    boolean isMatch = Pattern.matches("\b","\b");
    //isMatch is true 'cause "\b" in the first
    //argument is treated as a special character.
    //Same as the second argument.
    //
    //\b, as a special character, is a backspace
    System.out.println("isMatch? " + isMatch);
    
    //we escape the backslash by adding another
    //backslash before it. The additional
    //backslash overrides the special character
    //function of \b and changes its function to
    //a word boundary
    isMatch = Pattern.matches("\\b","\b");
    
    //isMatch is false 'cause "\\b" is treated
    //as a metacharacter. \b ,as a metacharacter,
    //is a word boundary
    System.out.println("isMatch? " + isMatch);
    
    //we add another backslash in this pattern,
    //the additional backslash overrides the
    //metacharacter function of \b and changes
    //its function back to a special character
    isMatch = Pattern.matches("\\\b","\b");
    //isMatch is true
    System.out.println("isMatch? " + isMatch);
    
    //we add another backslash in this pattern,
    //the additional backslash overrides the
    //special character function of \b and since
    //one of the backslashes already overrides
    //the metacharacter function of \b, the
    //matcher will treat backslash as a literal.
    isMatch = Pattern.matches("\\\\b","\\b");
    
    //isMatch is true 'cause "\\\\b" represents
    //"\b" literal, same as the second argument.
    System.out.println("isMatch? " + isMatch); 
    
    //This is how we put the literal backslash(\)
    //in a pattern
    isMatch = Pattern.matches("\\\\","\\");
    //isMatch is true
    System.out.println("isMatch? " + isMatch);
  }
}
Another way of escaping metacharacters is using the quote() method in Pattern class.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    
    //quote() method return a string object
    //where all characters are treated as
    //literals
    //
    //special meanings of characters like
    //metacharacters, special character, etc.
    //will be ignored
    boolean isMatch = Pattern.matches(Pattern.quote("[]"),"[]");
    //isMatch is true
    System.out.println("isMatch? " + isMatch);
  }
}
Groups and Capturing

Capturing group is a process of wrapping multiple characters into a single unit. If a capturing group gets a match against the input then, those matches will be saved in memory and can be recalled by using backreferences which we will discuss later.

Capturing groups are numbered from left to right. For example, we group our regular expression like this: ((A)(B(C)))
We can count the number of groups like this:
1. ((A)(B(C)))
2. (A)
3. (B(C))
4. (C)
To get how many groups are in an expression, use the groupCount() method in the matcher class. There's a group that is called "group 0". This group represents the whole expression and it's not counted as part of a capturing group.

There are two types of groups: capturing and non-capturing groups. The group that we discussed recently is a capturing group. Non-capturing groups are groups that don't capture string and not included in the total count of groups.

Groups that start with "(?" are either non-capturing groups or named-capturing groups. Embedded Flag Expressions are non-capturing groups that start with "(?".

Embedded Flag Expressions(Non-capturing groups)

Embedded Flag Expressions are non-capturing groups that act as an alternative to setting flags via compile() method with two arguments. We can include these expressions to set the flags that we need in an expression. These are the embedded flag expressions.
Pattern.CANON_EQ:  None
Pattern.CASE_INSENSITIVE:  (?i)
Pattern.COMMENTS:  (?x)
Pattern.MULTILINE:  (?m)
Pattern.DOTALL:  (?s)
Pattern.LITERAL:  None
Pattern.UNICODE_CASE:  (?u)
Pattern.UNIX_LINES:  (?d)

Let's create an example to demonstrate these expressions.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String input = "SixFive\nSixFour";
    //If we only need one flag then, we write
    //it like this: (?:X)pattern
    //where ?:X is the flag
    //e.g. "(?i)^six"
    String pattern = "(?i)(?m)^six";
    
    Matcher matcher = Pattern.compile(pattern).matcher(input);
    
    while(matcher.find()){
      System.out.println("Match Found!");
      System.out.print("start: " + matcher.start());
      System.out.print(" ");
      System.out.print("end: " + matcher.end());
      System.out.println();
    }
  }
}

Result
Match Found!
start: 0 end: 3
start: 8 end: 11
Capturing Groups

I already discussed the definition of capturing groups in the "Groups and Capturing" topic which can be seen above. Now, let's create an example to demonstrate capturing groups.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String pattern = "((http://|https://)(www\\.)"+
                     "(dailymotion|youtube))\\.com";
    String input = "https://www.youtube.com";
    
    Matcher matcher = Pattern.compile(pattern).
                      matcher(input);
    
    //Use groupCount() to count how many groups are in an expression
    System.out.println("Group Count: " + matcher.groupCount());
    
    if(matcher.lookingAt()){
      System.out.println("It's a match!");
      //input. Use the group(int index) method to get the captured
      //input that matched a portion of the input
      //the length and the max index count of capturing groups are 
      //equal
      //
      //Note: group() without argument returns the input
      //subsequence matched by the previous match
      System.out.println("Overall Match: " + matcher.group(0));
      System.out.println("1st group: " + matcher.group(1));
      System.out.println("2nd group: " + matcher.group(2));
      System.out.println("3rd group: " + matcher.group(3));
      System.out.println("4th group: " + matcher.group(4));
      
      //Use start(int group) and end(int group) to get the
      //start and end of the captured input in the expression.
      System.out.println();
      System.out.println("1st group start: " + matcher.start(1));
      System.out.println("3rd group start: " + matcher.start(3));
      System.out.println("4th group end: " + matcher.start(4));
    }
  }
}

Result
Group Count: 4
It's a match!
Overall Match: https://www.youtube.com
1st Group: https://www.youtube
2nd Group: https:
3rd Group: www.
4th Group: youtube

1st group start: 0
3rd group start: 8
4th group start: 12
Let's try another example.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    //The capture groups here are a bit complicated
    String pattern = "(.*?([Pp]([Aa])))\\d{2}";
    String input = "Peppa69";
    
    Matcher matcher = Pattern.compile(pattern).
                      matcher(input);
    
    System.out.println("Group Count: " + matcher.groupCount());
    
    if(matcher.matches()){
      System.out.println("It's a match!");
      
      for(int i = 0; i <= matcher.groupCount(); i++){
        System.out.println("Group " + i + ": " +
                           matcher.group(i));
      }
    }
    matcher.reset();
    
    System.out.println();
    while(matcher.find()){
      System.out.println("It's a match!");
      System.out.println("Group 2 start: " + matcher.start(2));
      System.out.println("Group 3 start: " + matcher.start(3));
    }
    
  }
}

Result
Group Count: 3
It's a match!
Group 0: Peppa69
Group 1: Peppa
Group 2: pa
Group 3: a

It's a match
Group 2 start: 3
Group 3 start: 4
The pattern in the above example is kinda complicated. My explanation about that complicated pattern is based on my intuition. So, take it with a grain of salt. I expect readers to have an understanding about the regex constructs that I used in the pattern above.

First off, the input is matched against the pattern. ".*?" grabs one character at the start of the input and then repeat the first and second steps until a match is found at index 3 where "p" that is near "a" is found. "p" and the subsequent "a" satisfy group 2 and 3. Group 2 match is found at index 3 and Group 3 match is found at index 4. Then, the subsequent numbers satify the "\\d++".

So, group0 is the combination of all subsequence that matches the patteren. Group1 is the combination of all subsequence in the group capturing that matches the pattern. Group2 is the combination of group2 and group3 subsequences. Group3 only has its own subsequence.

Backreferences

Backreference allows us to recall a pattern that is in a group. Backreference starts with a backslash, followed by a digit that represents the group number.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String pattern = "(\\d\\d)\\1";
    String input = "1212";
    
    Matcher matcher = Pattern.compile(pattern).
                      matcher(input);
    
    System.out.println("Group Count: " + 
                       matcher.groupCount());
    
    System.out.println();
    while(matcher.find()){
      System.out.println("Match found!");
      System.out.println("start: " + matcher.start());
      System.out.println("end: " + matcher.end());
    }
    
  }
}

Result
Group Count: 1

Match Found!
start: 0
end: 4
In the example above, we use "\1" to recall "(\d\d)" group. Try changing the input to "1234" and the matcher won't find any match. That's because, when we recall "\d\d", the backreference repeats the process that (\d\d) had done previously.

In previous process of (\d\d), it looked for two digits and found "1" and "2". Then, the backreference "\1" did the same thing. It looked for two digits and those digits must be "1" and "2". Let's try another example.
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String pattern = "([a-zA-Z]*?(\\d{3,3}))\\-\\2";
    String input = "Thud333-333";
    
    Matcher matcher = Pattern.compile(pattern).
                      matcher(input);
    
    System.out.println("Group Count: " + 
                       matcher.groupCount());
    
    System.out.println();
    while(matcher.find()){
      System.out.println("First Match");
      System.out.println("start: " + matcher.start());
      System.out.println("end: " + matcher.end());
      System.out.println();
      System.out.println("Group 0: " + matcher.group(0));
      System.out.println("Group 1: " + matcher.group(1));
      System.out.println("Group 2: " + matcher.group(2));
      System.out.println();
    }
    
    pattern = "([a-zA-Z]*?(\\d{3,3}))\\-\\1";
    input = "Thud333-333";
    
    matcher = Pattern.compile(pattern).
              matcher(input);
              
    System.out.println();
    while(matcher.find()){
      System.out.println("Second Match");
      System.out.println("start: " + matcher.start());
      System.out.println("end: " + matcher.end());
      System.out.println();
      System.out.println("Group 0: " + matcher.group(0));
      System.out.println("Group 1: " + matcher.group(1));
      System.out.println("Group 2: " + matcher.group(2));
      System.out.println();
    }
    
  }
}

Result
Group Count: 2

First Match
start: 0
end: 11

Group 0: Thud333-333
Group 1: Thud333
Group 2: 333

Second Match
start: 0
end: 11

Group 0: 333-333
Group 1: 333
Group 2: 333
So, in the first match, the most suitable match for "([a-zA-Z]*?(\\d{3,3}))\\-\\2" pattern is "Thud333-333" or the entire input. "\2" is a backreference that refers to "(\\d{3,3})" group.

In the second match, the most suitable match for "([a-zA-Z]*?(\\d{3,3}))\\-\\1" is "333-333". "\1" is a backreference the refers to "([a-zA-Z]*?(\\d{3,3}))" group.

Named Capturing Groups

We can put names in our capturing groups. Capturing groups are still numbered even we assign names in them.
Syntax: (?<name>X)
import java.util.regex.*;
public class SampleClass{

  public static void main(String[]args){
    String pattern = "(?<num>(\\d){4})\\-\\1";
    String input = "1234-1234";
    
    Matcher matcher = Pattern.compile(pattern).
                      matcher(input);
    
    System.out.println("Group Count: " + 
                       matcher.groupCount());
    
    System.out.println();
    while(matcher.find()){
      System.out.println("Match found!");
      System.out.println("start: " + matcher.start());
      System.out.println("end: " + matcher.end());
      System.out.println();
      
      System.out.println("Group 0: " + matcher.group(0));
      //we use the name of group1 to get the subsequence
      //in it
      System.out.println("num: " + matcher.group("num"));
      System.out.println("Group 2: " + matcher.group(2));
    }
    
  }
}

Monday, July 5, 2021

Java Tutorial: Annotations

Chapters

Java Annotation

Annotations were introduced in java5 and their purpose is to provide metadata for our code. Annotations can also be used to set instructions to the compiler; build instructors for compile-time that can be used for generating code like XML; create metadata and instructions for our code that can be available during runtime.

Annotations don't directly affect the behavior of our code. In other words, They don't directly affect the functionality of our program. Annotations can be applied to methods, fields, classes and interfaces.

To annotate an element, write the annotation followed by the element.
e.g.
Annotation syntax
without value
@annotation-name
with value
@annotation-name("value")

Annotating fields/variables
@Deprecated int deprecation = 10;
or
@Deprecated
int deprecation = 10;

Annotating blocks(e.g. methods, constructors, etc.)
@Deprecated public void meth(){/**/}
or
@Deprecated
public void meth(){/**/}

Multiple annotations can be in any order
@FunctionalInterface
@Deprecated
interface InterfaceA{
  void meth();
}
or
@Deprecated
@FunctionalInterface
interface InterfaceA{
  void meth();
}

Built-in Annotations

Java has some built-in annotations. These are: @Deprecated, @FunctionalInterface, @Override, @SafeVarargs and @SuppressWarnings. These annotations are ready to use and give helpful instructions to the compiler.

@Deprecated

A program element(e.g. field,method,etc.) that is annotated as @Deprecated is an element that programmers are discouraged from using. An element may be deprecated if it's obsolete or a new version/alternative is implemented. Using deprecated element may cause errors, incompatibility, etc.

When we use a pre-deprecated element annotated with @Deprecated, The compiler will give us a warning.
public class SampleClass{

  public static void main(String[]args){
    
    //Integer constructor is deprecated
    //in java9 and above. The compiler
    //will give us warning if we use
    //Integer constructor
    Integer intOne = new Integer(10);
    
    //Integer.valueOf() is the Recommended
    //alternative by java for deprecated
    //Integer constructor with int parameter
    //Integer intOne = Integer.valueOf(10);
  }
}
We can use @Deprecated if we want a particular element of our code to be deprecated.
public class SampleClass{
  
  //make this method obsolete
  //by using @Deprecated annotation
  //
  //This implementation doesn't
  //consider the possibility of
  //two numbers being equal.
  //
  //@Deprecated can have value
  //like this: @Deprecated(forRemoval=true)
  //We will discuss annotation
  //values later
  //
  @Deprecated
  int findMax(int n1, int n2){
    
    if(n1 > n2)
      return n1;
    else return n2;
    
  }
  
   //new and improved version
   Integer findMaxInteger(Integer n1, Integer n2){
    Integer result = null;
    
    if(n1 > n2)
      result = n1;
    else if(n1 < n2)
      result = n2;
      
    return result;
  }
  
  public static void main(String[]args){
    SampleClass sc = new SampleClass();
    
    //unlike pre-deprecated elements in
    //java, our user-deprecated elements
    //won't make the compiler to throw 
    //a warning regarding deprecation
    //
    //I'm not sure if this is compiler-specific
    //or not.
    //
    //As the time of this writing, I'm using
    //JDK11 of AdoptOpenJDK and my compiler
    //is not giving any warning regarding
    //user-deprecated elements
    //int myInt = findMax(10,15);
    
    Integer intOne = sc.findMaxInteger(10,15);
    if(intOne != null)
      System.out.println("Max: " + intOne);
    else System.out.println("intOne is null!");
    
  }
}
You might ask: "Why don't we remove the obsolete method instead?". Well, It's a bad idea if we do that. Imagine you created a library and some developers are using your created library. Then, You upgraded the version of your library and remove findMax() instead of depracating it. Now, the developers wanna use your library's new version.

Let's say they use your library's new version. Since you remove findMax(), the developers can't compile their codebase anymore because findMax() is missing. If the program is a huge one then their migration is a disaster and most likely, they will rollback to the previous version.

So, it's better to inform your users about the obsolete method first than removing it without any notification. This way, we're giving them time to adapt to your new version.

@FunctionalInterface

We use @FunctionalInterface to give an instruction to the compiler that an interface is a functional interface and the interface must follow the functional interface rule where only one abstract method is allowed.
public class SampleClass{

  public static void main(String[]args){
  }
}

@FunctionalInterface
interface ConcatString{

  void concat(String s1, String s2);
  
  //compiler will throw a compile-time
  //error 'cause we violate the 
  //functional interface rule if we 
  //another abstract method
  //void concatString(String[] s);
}
If we uncomment concatString() method, the compiler will throw a compile-time error. If we uncomment concatSting() and remove @FunctionalInterface then, the compiler won'y throw a warning regarding functional interface even ConcatString interface violates the functional interface rule.

If you intend an interface to be a functional interface then, putting @FunctionalInterface on top of the interface may be helpful.

@Override

Note: You need to have an understanding of method overriding to understand the usage of @Override.
We can use the @Override to give instructions to the compiler that a method annotated with @Override overrides a method in superclass.
public class SampleClass{
  
  public static void main(String[]args){
    ClassB b1 = new ClassB("SampleClass");
  }
}

class ClassA{

  void meth(String s){
    System.out.println("ClassA: " + s);
  }
}

class ClassB extends ClassA{

  @Override
  void meth(String s){
    System.out.println("ClassB: " + s);
  }
}
This example above will compile just fine. @Override ensures that we're overridng a method in a superclass. Try removing the meth() in ClassA or change the method signature of the overriding method and the compiler will throw a compile-time error.

@SafeVarargs

@SafeVarargs is solely used to suppress a warning that is related to variable arguments, which is the heap pollution warning. Use @SafeVarargs if you're confident that your generic varargs is safe to use.
public class SampleClass{
  
  //This method won't issue a warning regarding
  //heap pollution
  @SafeVarargs
  static <T>void displayString(T... obj){
    for(Object o : obj)
      System.out.println(o.toString());
  }
  
  public static void main(String[]args){
    SampleClass.displayString();
  }
}
Try removing @SafeVarargs and recompile the example above. This time, the compiler will throw a warning.

@SuppressWarnings

As the annotation name implies, @SuppressWarnings is used to suppress warnings. Unlike other built-in annotations, this annotation requires a value in array form. This annotation can suppress various warnings like deprecation, unchecked, etc.
General Form: @SuppressWarnings({"warning-to-suppress"})

public class SampleClass{
  
  public static void main(String[]args){
    
    //This instantiation won't issue a
    //deprecation warning
    //
    //Tip: Don't suppress deprecation
    //warning just because you want
    //to. Deprecated element may cause
    //errors so be careful using them.
    //
    //You may wanna suppress deprecation
    //if you're doing a test for 
    //the deprecated element
    @SuppressWarnings({"deprecation"})
    Integer intOne = new Integer(10);
    
    //suppressing raw types warning.
    //raw types warning is part of 
    //unchecked warnings
    //
    //Note: in modern java application,
    //raw types should be avoided
    @SuppressWarnings({"unchecked"})
    ClassA<T> classA = new ClassA("name");
  }
}

class ClassA<T>{
   T name;
   
   ClassA(T name){
     this.name = name;
   }
}
When we suppress a warning at a class level then, its members that throw the same warning will also be suppressed.
//This suppressWarnings annotation suppresses two warnings:
//
//"varargs" and "unchecked".
//"varargs" warning is issued when use generic varargs.
//Generic varargs cause heap pollution.
//
//"unchecked" warning is issued when the 
//compiler can't fully perform type checking
//Java doesn't allow us to create arrays of
//parameterized types. Consequently, elements
//of an array can't be generic. thus,
//the compiler can't fully peform type
//checking on arrays.
//varargs is considered as an array
@SuppressWarnings({"varargs","unchecked"})
public class SampleClass{
  
  //No need to put @SafeVarargs here. @SafeVarargs
  //is equivalent to @SuppressWarnings({"varargs","unchecked"})
  //@SafeVarargs
  static <T>void displayString(T... obj){
    for(Object o : obj)
      System.out.println(o.toString());
  }
  
  public static void main(String[]args){
    SampleClass.displayString();
  }
}
Creating Custom Annotation

We can create our custom annotation by using the @interface syntax. Don't be confused between @interface and interface keyword. @interface is used for custom annotation whereas interface keyword is used for creating interfaces.
General Form

access-modifier @interface annotation-name{

  //annotation body
}
We can put elements/values in our annotation and they look like methods. However, their functionalities are way different from methods and we shouldn't add implementation to these method-like elements in annotation. The purpose of annotation element is to store a value like string or number.

Also, all annotations including our custom annotations implicitly extend java.lang.annotation.Annotation interface and java doesn't allow annotations to explicitly extend any java class/interface/annotation.
public class SampleClass{
   //Annotation elements without default values
   //requires to be initialized when the annotation
   //where they reside is used
   //this are the formats for assigning values to
   //elements: 
   //element-name="value" for string element
   //element-name=number e.g. version=1.555.555 for number element
   //element-name{"value1",number1,"value2"} for array element
   //element-name{"value1"} for array element with single value
   @Author(authorName="Ghosts",authorEmail="MyAddress@yahoo.com")
   @Contributors(names={"Dan Brooks","David Muller"})
   //initializing values in annotation elements with default values
   //is optional. As we can see, @ComparisonAlgorithm annotation
   //elements are not required to be initialized
   @ComparisonAlgorithm
   Integer findMaxInteger(Integer n1, Integer n2){
    Integer result = null;
    
    if(n1 > n2)
      result = n1;
    else if(n1 < n2)
      result = n2;
      
    return result;
  }
  
  //initializing annotation elements
  //with default values is optional
  @ComparisonAlgorithm(dataType="Primitive",returnNull=false)
  int findMax(int n1, int n2){
    
    if(n1 > n2)
      return n1;
    else return n2;
    
  }
  
  public static void main(String[]args){
  }
  
  //custom annotation can be nested
  //just like regular interface
  @interface Author{
  
    //annotation elements
    //without default value
    String authorName();
    String authorEmail();
    
    //annotation elements
    //with default value
    //we will use the default
    //keyword to specify 
    //the default value of
    //an element
    //String authorName() default "Brainy Ghosts";
    //String authorEmail() default "SampleAddress@yahoo.com";
    
  }
}

@interface Contributors{
  
  //an array annotation
  //element without default
  //values
  String[] names();
  
  //an array annotation
  //element with default
  //values
  //String[] names() default {"Dan Brooks","David Muller"};
  
}

@interface ComparisonAlgorithm{

  int version() default 1;
  String dataType() default "Object";
  boolean returnNull() default true;
}
If our annotation has only one element, we can name that element as "value" then, we are not required to include the element name of that element if we initialize the annotation.
public class SampleClass{
   
   //We don't need to include element name
   //If each annotation has a single element
   //named as "value"
   @Author("Ghosts")
   @Contributors({"Dan Brooks","David Muller"})
   Integer findMaxInteger(Integer n1, Integer n2){
    Integer result = null;
    
    if(n1 > n2)
      result = n1;
    else if(n1 < n2)
      result = n2;
      
    return result;
  }
  
  public static void main(String[]args){
  }
}

@interface Author{
  String value();  
}

@interface Contributors{
  String[] value();
}
One of the usage of annotation is to provide metadata for our code. As we can see in the example above, we provide additional information for findMaxInteger() and findMax() methods. If we want to access annotation during runtime, we need to add @Retention annotation on our annotation block then, we can use java reflection to access annotation and its elements.

Annotations for Creating Annotation

Java has annotations that can be used to increase or reduce restrictions of declared annotation. These annotations are different from the built-in annotations. The annotations that we're going to discuss here is solely used for annotation type declaration/definition whereas built-in annotations are used for java elements like constructor,method,field, etc.

Though, There are built-in annotations that can be used for annotation declaration like @Deprecated. To use annotations used for annotation declaration, we need to import java.lang.annotation package.

@Target

@Target specifies where annotations can be used. @Target accepts ElementType as values. If we don't define @Target in our annotation declaration then, that annotation can be used in any java element.
import java.lang.annotation.*;
//valid
@Author(name="Ghosts")
//invalid
//@ComparisonAlgorithm
public class SampleClass{

  //valid
  @ComparisonAlgorithm
  Integer findMaxInteger(Integer n1, Integer n2){
    Integer result = null;
    
    if(n1 > n2)
      result = n1;
    else if(n1 < n2)
      result = n2;
      
    return result;
  }

  public static void main(String[]args){
  }
}

//This annotation doesn't have
//@Target annotation. Thus,
//this annotation can be used
//in any java element
@interface Author{
  String name();
}

//This annotation has @Target
//annotation. This annotation
//can only be used in method
//elements
//
//We can add multiple element
//types in @Target by separating
//them using comma(,)
//e.g. @Target({ElementType.METHOD,ElementType.FIELD})
//For more information about @Target
//visit the java documentation
@Target({ElementType.METHOD})
@interface ComparisonAlgorithm{

  int version() default 1;
  String dataType() default "Object";
  boolean returnNull() default true;
}
@Documented

If we're creating a documentation about our codebase and we want our annotations to be included in that documentation then, We need to annotate our annotation with @Documented annotation.
import java.lang.annotation.*;

@Author(name="Ghosts")
public class SampleClass{

Integer findMaxInteger(Integer n1, Integer n2){
  Integer result = null;
    
  if(n1 > n2)
    result = n1;
  else if(n1 < n2)
    result = n2;
      
  return result;
  }

  public static void main(String[]args){
  }
}

//This annotation will be included in
//a documentation, if we create one.
//Use javadoc tools to create a 
//documentation of your codebase
@Documented
@interface Author{
  String name();
}
@Retention

@Retention specifies the retention of annotations. @Retention accepts RetentionPolicy as value. There are three RetentionPolicy values that we can use. These are: RetentionPolicy.SOURCE, RetentionPolicy.CLASS and RetentionPolicy.RUNTIME.

RetentionPolicy.SOURCE means that annotations is only available in the source code.

RetentionPolicy.CLASS means that annotations is included in the .class file. Users can view annotations if they inspect .class file.

RetentionPolicy.RUNTIME means that annotations can be accessed during runtime. Use java reflection tools to access annotations and their elements.

If we don't annotate annotations with @Retention then, the RetentionPolicy value of those annotations is going to be RetentionPolicy.CLASS by default.
import java.lang.reflect.*;
import java.lang.annotation.*;

public class SampleClass{
   @Author(authorName="Ghosts",authorEmail="MyAddress@yahoo.com")
   @Contributors(names={"Dan Brooks","David Muller"})
   @ComparisonAlgorithm
   Integer findMaxInteger(Integer n1, Integer n2){
    Integer result = null;
    
    if(n1 > n2)
      result = n1;
    else if(n1 < n2)
      result = n2;
      
    return result;
  }
  
  //initializing annotation elements
  //with default values is optional
  @ComparisonAlgorithm(dataType="Primitive",returnNull=false)
  int findMax(int n1, int n2){
    
    if(n1 > n2)
      return n1;
    else return n2;
    
  }
  
  //display Annotation elements using reflection tools
  static void displayElements(Annotation[] annotations){
    for(Annotation anno : annotations){
      if(anno instanceof Author){
        Author author = (Author)anno;
        System.out.println(author.authorName());
        System.out.println(author.authorEmail());
      }
      else if(anno instanceof ComparisonAlgorithm){
        ComparisonAlgorithm ca = (ComparisonAlgorithm)anno;
        System.out.println(ca.version());
        System.out.println(ca.dataType());
        System.out.println(ca.returnNull());
      }
      //this block won't execute 'cause the 
      //Contributors annotation can't be accessed
      //at run-time
      else if(anno instanceof Contributors){
        Contributors cb = (Contributors)anno;
        for(String s : cb.names())
          System.out.println(s);
      }
      System.out.println();  
    }
    
  }
  
  public static void main(String[]args){
    //we're going to use java reflection tools
    //here to inspect annotations at runtime
    Class sc = SampleClass.class;
    Method[] meth = sc.getDeclaredMethods();
    
    for(Method m : meth)
      SampleClass.displayElements(m.getAnnotations());
      
  }
  
  //This annotation can be accessed at 
  //runtime
  @Retention(RetentionPolicy.RUNTIME)
  @interface Author{
    String authorName();
    String authorEmail(); 
  }
}

//This annotation has 
//RetentionPolicy.Class by default.
//Thus, It can't be accessed at
//runtime
@interface Contributors{
  String[] names();
}

//This annotation can be accessed at 
//runtime
@Retention(RetentionPolicy.RUNTIME)
@interface ComparisonAlgorithm{

  int version() default 1;
  String dataType() default "Object";
  boolean returnNull() default true;
}
@Inherited

If an annotation with @Inherited annotates an element and that element has subclasses then, the subclasses will be annotated with the same annotation of their superclass.
import java.lang.reflect.*;
import java.lang.annotation.*;
public class SampleClass{

  public static void main(String[] args){
    Class b1 = ClassB.class;
    
    //ClassB inherited the Author annotation of
    //ClassA. Try removing @Inherited in Author
    //declaration and the length of this array
    //will be zero
    Annotation[] annotations = b1.getAnnotations();
    System.out.println("# of Annotations: " + annotations.length);
    
    for(Annotation anno : annotations)
      if(anno instanceof Author){
        Author author = (Author)anno;
        System.out.println(author.value());
      }
        
  }
}

@Author("Brainy Ghosts")
class ClassA{
}

class ClassB extends ClassA{
}

@Retention(RetentionPolicy.RUNTIME)
@Inherited
@interface Author{
  String value();
}
According to @Inherited Documentation: "this meta-annotation type has no effect if the annotated type is used to annotate anything other than a class. Note also that this meta-annotation only causes annotations to be inherited from superclasses; annotations on implemented interfaces have no effect".

Saturday, July 3, 2021

Java Tutorial: Wrapper Classes, Autoboxing and Auto-unboxing

Chapters

Java Wrapper Classes

Note: Some parts of this tutorial is going to use generics. So, an understanding of generics is required at some parts of this tutorial.

Wrapper classes are object wrappers for primitive types. These are wrapper classes in java: Boolan, Byte, Character, Float, Integer, Long, Short, Double. We use wrapper classes to wrap primitive types into object types so that, we can pass them to other objects that don't accept primitive types like List<>, ArrayList<>, etc.

ArrayList can't have primitive types as elements, we need to wrap primitive types first before ArrayList can accept them as its elements. Take a look at this example.
import java.util.ArrayList;
public class SampleClass{

  public static void main(String[]args){
    ArrayList<Integer> arrList = new ArrayList<>();
    arrList.add(1);
    
    Sytem.out.println(arrList.get(0));
  }
}
The code above will run just fine. You might say: "But you say that primitive types can't be elements of arraylist?". Well, my statement is still correct this add() method arrList.add(1); accepts Integer as argument. The primitive "1" is automatically converted to Integer.

Autoboxing

Autoboxing is the automatic conversion of primitives to their respective wrapper classes. The example above demonstrates autoboxing. Regardless, let's create another example.
public class SampleClass{

  public static void main(String[]args){
    //autoboxing
    Integer intOne = 1;
    
    //manual boxing
    Integer intTwo = Integer.valueOf(2);
  }
}
Imagine calling Integer.valueOf() everytime we assign an int primitive to Integer. Autoboxing removes the use of manual boxing if we're just assigning a primitive to its respective wrapper class. Take a look at this example.
public class SampleClass{

  public static void main(String[]args){
  
    //error: String can't be converted
    //to Integer
    //Integer intOne = "1";
    
    //If we want a String to be converted
    //to Integer then, we need to 
    //manually convert the String
    Integer intTwo = Integer.valueOf("1");
  }
}
Auto-unboxing

Auto-unboxing is the automatic conversion of wrapper classes to their respective primitive types. Take a look at this example.
public class SampleClass{

  public static void main(String[]args){
    //autoboxing
    Integer intOne = 1;
    
    //auto-unboxing
    int myInt = intOne;
    
    //manual unboxing
    //int myInt = intOne.intValue();
    
    System.out.println(myInt);
  }
}
Auto-unboxing removes the use of manual unboxing if we're just assigning a wrapper object to its respective primitive type.

Using Operators on Wrapper Classes

Now we know autoboxing and auto-unboxing, we also know the reason why we can use operators on wrapper classes. When an operator is used on wrapper classes, Wrapper classes are automatically converted to their respective primitive types first before performing the operation.

The result may be converted depending on the type of the variable where the result is stored.
public class SampleClass{

  public static void main(String[]args){
    Integer a1 = 1;
    Integer a2 = 3;
    
    //a1 and a2 are unboxed first
    //before performing the operation
    //then assign the result
    int intOne = a1 + a2;
    System.out.println(intOne);
    
    //a1 is unboxed first
    //before performing the operation
    //then assign the result
    intOne = a1 + 2;
    System.out.println(intOne);
    
    //the operation is performed
    //then the primitive result is
    //wrapped into Integer and
    //then it's assigned to a3
    Integer a3 = 1 + 4;
    System.out.println(a3);
    
    //a1 is unboxed first
    //before performing the operation
    //then the result is boxed
    //before it's assigned to a3
    a3 = a1 + 1;
    System.out.println(a3);
    
    //a1 and a2 are unboxed first
    //before performing the operation
    //then the result is boxed 
    //before it's assigned
    a3 = a1 + a2;
    System.out.println(a3);
  }
}
Now, let's create the manual box/unbox version of the example above.
public class SampleClass{

  public static void main(String[]args){
    Integer a1 = 1;
    Integer a2 = 3;
    
    int intOne = a1.intValue() + a2.intValue();
    System.out.println(intOne);
    
    intOne = a1.intValue() + 2;
    System.out.println(intOne);
    
    Integer a3 = Integer.valueOf(1 + 4);
    System.out.println(a3);
    
    a3 = Integer.valueOf(a1.intValue() + 1);
    System.out.println(a3);
    
    a3 = Integer.valueOf(a1.intValue() + a2.intValue());
    System.out.println(a3);
  }
}
When performing operation between primitive types, typecasting(Primitive Types) may happen.
public class SampleClass{

  public static void main(String[]args){
    
    //error: float can't be converted
    //to Integer
    //Integer intOne = 30.5f + 10;
    
    //valid
    Float floatOne = 30.5f + 10;
    System.out.println(floatOne);
    
    //error: int can't be converted to
    //Short
    //short st = 50;
    //Short shortOne = st + 50;
  }
}

Friday, July 2, 2021

Java Tutorial: Lambda Expression

Chapters

Lambda Expression

Lambda Expression was introduced in java8. It's a clear and compact way of defining a function of an abstract method in a functional interface. Functional interface is an interface with one abstract method only. Let's create an example demonstrating lambda expression.
General Form: (arg1,arg2) -> {};
Where:
(arg1,arg2) - argument-list
-> - arrow token
{}; - body


public class SampleClass{
  
  static void displayString(PrintString ps,
                            String s){
      if(s.startsWith("My"))
         ps.printString(s);
    
  }
  
  public static void main(String[]args){
    
    //Standard way of writing lambda.
    //
    //we can change the argument name in the argument-list
    //even the name doesn't match with the
    //parameter name of functional interface
    //and this code will still run fine.
    //
    //for convenience, we want the argument-list of
    //lambda to have the same name as the
    //parameter-list of the abstract method
    //in functional interface.
    SampleClass.displayString((String s) -> 
                             {System.out.println(s);},
                             "MyString1");'
                             
    //Lambda body({}) can be omitted if the body
    //only has a single method call statement.
    //Some statements require the body like the return
    //statement,assignment statement, etc.
    //e.g
    //() -> return true; //invalid
    //() -> {return true;}; //valid
    //() -> System.out.println(); //valid
    //
    //Types in argument-list can also
    //be omitted.
    //e.g. (String s) = (s)
    //Sometimes, java can't infer the type from the
    //parameter-list. Java will notify us about this
    //and if we receive a message then, we need
    //to explicitly write the type.
    //Note that if one element in the argument-list
    //has or doesn't have type then, all elements
    //in the argument-list must have or mustn't have
    //types
    //
    //We can omit the parentheses of argument-list
    //if there's only one argument in the argument-list
    //e.g (s) = s
    //
    //We can store lambda expression in a
    //functional interface variable.
    PrintString ps = s -> System.out.println(s);
    SampleClass.displayString(ps,"MyString2");
    
    //displayString() lambda with omitted body({})
    /*
        SampleClass.displayString((String s) -> 
                                  System.out.println(s),
                                  "MyString1");
    */
  }
}

@FunctionalInterface
interface PrintString{
  void printString(String s);
}
Now, let's try lambda expression with return type and no argument-list.
public class SampleClass{

  public static void main(String[]args){
    
    //leave the argument-list blank
    //since sayHello() abstract method in SayHello
    //interface doesn't have parameters
    SayHello sh = () -> System.out.println("Hello");
    sh.sayHello();
    
    //We can put multiple statements in lambda's
    //body
    CompareStringLength compStrLength = 
    (s1,s2) -> { 
      //if s1 is greater than s2
      if(s1.length() > s2.length())
         return 1;
      //if s2 is greater than s1
      else if(s1.length() < s2.length())
         return -1;
      
      //If both strings have equal length
      return 0;
    };
    
    switch(compStrLength.compareLength("String1","String2")){
    
      case 1:
      System.out.println("s1 is greater than s2");
      break;
      
      case -1:
      System.out.println("s2 is greater than s1");
      break;
      
      case 0:
      System.out.println("s1 and s2 have equal length");
      break;
    }
    
    //if the lambda body only has a return statement
    //we can omit the return keyword like in this
    //lambda expression
    StringEquality se =
    (s1,s2) -> s1.equals(s2);
    
    //the stament above is equivalent to this
    //statement
    //StringEquality se =
    //(s1,s2) -> {return s1.equals(s2);};
    
    boolean result = se.checkEquality("A","A");
    System.out.println(result);
  }
}

@FunctionalInterface
interface SayHello{
  //abstract method with no
  //parameters and return
  //type
  void sayHello();
}

@FunctionalInterface
interface CompareStringLength{
  
  //abstract method with parameters
  //and return type.
  int compareLength(String s1, String s2);
}

@FunctionalInterface
interface StringEquality{

  boolean checkEquality(String s1, String s2);
}
var Type Name in Lambda Argument-List

Starting from java11, We can use the var Type Name as type in lambda argument-list.
public class SampleClass{
  
  static void displayString(PrintString ps,
                            String s){
      if(s.startsWith("My"))
         ps.printString(s);
    
  }
  
  public static void main(String[]args){
    //use var keyword as type on argument-list
    PrintString ps1 = (var s) -> System.out.println("String:" + s);
    
    SampleClass.displayString(ps1,"MyString1");
  }
}

@FunctionalInterface
interface PrintString{
  void printString(String s);
}
Lambda Variable Capture

Lambda expression can capture variables outside of its scope. Lambda expression can capture local, instance and static variables.
public class SampleClass{
  //instance variable
  String str1 = "str1";
  //static variable
  static String str2 = "str2";
  
  void displayString(){
    //local variable
    //Note: local variables referenced
    //from a lambda expression must be
    //final or effectively final
    String str3 = "str3";
    
    //this lambda expression captures
    //str1,str2 and str3
    ConcatString cs = () ->
    {return str1 + str2 + str3;};
    
    System.out.println(cs.concatString());
  }
  
  public static void main(String[]args){
    SampleClass sc = new SampleClass();
    sc.displayString();
  }
}

@FunctionalInterface
interface ConcatString{

  String concatString();
}
Be careful capturing global variables in multithreading. Even though global variables can be captured even they're not final or effectively final, the values they're holding may not change if they're updated in the lambda expression. Take a look at this example.

Note: This example may cause infinite loop, you should know how to close JVM by force before testing this example.
public class SampleClass{
  
  volatile static int count = 0;
  public static void main(String[] args){
    
    while(count < 3){
      System.out.println("To infinity and beyond...");
      new Thread(() -> {
        count++;
      });
    }
  }
}
Method Reference as Lambda Expression

We can substitute a method/constructor call as lambda expression. The parameters, parameter type and return type of the method definition of that method call should match the abstract method parameters, parameter type and return type in the functional interface.

First, we need to convert the method/constructor call as method reference. Method reference is similar to method call structure but without the parentheses and with additional "::" symbol.
Syntax: class-name::method-name;
public class SampleClass{

  public static void main(String[]args){
    //println is a member of out class in
    //System class. So, we refer to the System
    //then refer to out and refer to println().
    DisplayString ds = System.out::println;
    //The statement above is equivalent to this lambda
    //DisplayString ds = s -> System.out.println(s);
    
    ds.displayString("String");
  }
}

@FunctionalInterface
interface DisplayString{

  void displayString(String s);
}
In the example above, we see that the functionality of println() becomes the functionality of displayString() in DisplayString interface. We can also create method references from the available methods of a parameter of the abstract method in functional interface.
public class SampleClass{

  public static void main(String[]args){

    StringCaps sc = String::toUpperCase;
    
    String str = sc.capitalizeString("String");
    
    System.out.println(str);
  }
}

@FunctionalInterface
interface StringCaps{

  String capitalizeString(String s);
}

Result
STRING
In the example above, we create a method reference of String.toUpperCase(). This is possible because the parameter is capitalizeString() is a String type and toUpperCase() is available to any String instance.

Method Reference with Unequal Parameter Count

There are methods that can be referred to a functional interface even the parameter count of those method don't match the parameter of the abstract method in functional interface. The first parameter in the functional interface calls the referenced method the others are used as arguments.
public class SampleClass{

  public static void main(String[]args){
    
    //concat() in String class only has
    //one parameter and concat() in
    //ConcatString has two parameters.
    ConcatString cs = String::concat;
    
    //The statement above is equivalent to 
    //this lambda.
    //
    //As we can see, s1 calls concat() of String
    //class and s2 is the argument for concat() of
    //String.
    //ConcatString cs = (s1,s2) -> s1.concat(s2);
    
    System.out.println(cs.concat("A","B"));
  }
}

@FunctionalInterface
interface ConcatString{

  String concat(String s1,String s2);
}
Try changing the parameter of concat() in ConcatString and you will get an invalid method reference error.

Method Reference Return Type and Parameter Type

It's alright if the return type of the method in functional interface and the return type of the referenced method are not the same, as long as those return types are related. For parameters, the parameter type of the method in functional interface should be equal.
public class SampleClass{

  public static void main(String[]args){
    
    ConcatString cs = String::concat;
    
    System.out.println(cs.concat("A","B"));
  }
}

@FunctionalInterface
interface ConcatString{
  
  //the return type of concat() in String class
  //is String. CharSequence and String are
  //related, so concat() of String can be
  //referenced to this method
  CharSequence concat(String s1,String s2);
  
  //We can't reference concat() of String to this
  //method 'cause the first parameter is CharSequence and
  //CharSequence doesn't have concat(). So, s1 can't call
  //concat()
  //
  //Second, the second parameter is CharSequence and the
  //parameter of concat() of String class is String.
  //We can't downcast a parameter type. So, the second
  //parameter is invalid.
  //CharSequence concat(CharSequence s1,CharSequence s2);
}
We can also refer a functional interface with void return type to a method reference with a return type that is not void. However, we can't refer a functional interface with non-void return type to a method reference with void return type.
public class SampleClass{

  public static void main(String[] args){
    
    //invalid
    //interface1 i1 = SampleClass::meth;
    
    //valid
    interface2 i2 = SampleClass::meth2;
    i2.invoke(2);
  }
  
  static void meth(int num){
    System.out.println(num*num);
  }
  
  static int meth2(int num){
    System.out.println(num*num);
    return num*num;
  }
}

@FunctionalInterface
interface interface1{

  int invoke(int num);
}

@FunctionalInterface
interface interface2{

  void invoke(int num);
}
User-Defined Methods as Method Reference

We can also use our own method as method reference. Static and non-static methods can be referred.
public class SampleClass{
  private String name;
  
  SampleClass(String name){
    this.name = name;
  }
  
  static void displayString(CharSequence s){
    System.out.println(s);
  }
  
  String concatName(SampleClass sc){
    return name.concat(sc.toString());
  }
  
  @Override
  public String toString(){
    return name;
  }
  
  boolean checkEquality(String s1, String s2){
    return s1.equals(s2);
  }
  
  public static void main(String[]args){
    //using static method as method reference
    //also, notice the parameter of displayString()
    //and printString().
    //printString() has a parameter type of String
    //whereas displayString() has a parameter type
    //of CharSequence.
    //String can be upcasted to CharSequence.
    //That's why this method reference is valid.
    PrintString ps = SampleClass::displayString;
    ps.printString("String");
    
    SampleClass sc1 = new SampleClass("SampleClass1");
    
    //using non-static method as method reference
    //using class name
    CombineString cs = SampleClass::concatName;
    
    String str = cs.concat(sc1,new SampleClass("SampleClass2"));
    System.out.println(str);
    
    //using non-static method as method reference
    //using variable name
    CompareString compStr = sc1::checkEquality;
    System.out.println(compStr.compare("Str1","str2"));
    
    //invalid: when we refer a method to a functional interface
    //using class instance, we must match the parameter count
    //of the functional interface to the parameter count of
    //referenced method.
    //concatName has one parameter whereas concat() in
    //CombineString has two parameters. So, this
    //method reference is invalid
    //CombineString cs = sc1::concatName;
  }
}

@FunctionalInterface
interface PrintString{

  void printString(String s);
}

@FunctionalInterface
interface CombineString{

  String concat(SampleClass sc1, SampleClass sc2);
}

@FunctionalInterface
interface CompareString{

  boolean compare(String s1,String s2);
}
Constructor Reference

We can refer a constructor to a functional interface. We need to tweak our method reference syntax to be suitable for constructors. This is our new syntax: class-name::new;
public class SampleClass{

  public static void main(String[]args){
    //using constructor as a reference for
    //StringInstance
    StringInstance si = String::new;
    //The statement above is equivalent to
    //this lambda
    //StringInstance si = letters -> new String(letters);
    
    String s1 = si.createString(new char[]{'s','t'});
    String s2 = si.createString(new char[]{'r','i','n','g'});
    
    System.out.println(s1.concat(s2));
  }
}

@FunctionalInterface
interface StringInstance{

  String createString(char[] letters);
}
The parameters of the method in functional interface must match the parameters of constructor. Also, we can see that createString() method is similarly behaving like a constructor.
java.util.function Introduction

Note: You need to have an understanding about generics to understand generics syntax that we're going to discuss here. java.util.function package has lots of pre-built functional interfaces that we can use. As the title of this topic implies, I'll only introduce basic usage of some functional interfaces in this package.

Let's try the Predicate<T> interface.
import java.util.function.Predicate;
public class SampleClass{

  public static void main(String[]args){
    String str = "myString";
    
    //Predicate<T> is a functional
    //interface that is part of
    //java.util.function package
    Predicate<String> predicate =
    (t) -> t.toString().equals(str);
    
    //test() method in Predicate accepts "T"
    //type as argument and return a boolean value
    boolean result = predicate.test("MyString");
    System.out.println(result);
  }
}
Next, let's try the Consumer<T> interface.
import java.util.function.Consumer;
public class SampleClass{
  static String str = "string";
  
  public static void main(String[]args){
    
    Consumer<String> consumer = 
    t -> str = str.concat(t.toString());
    
    Consumer<CharSequence> after =
    t -> System.out.println(t.toString() + str);
    
    //andThen(Consumer<? super T> after)
    //returns a composed Consumer that
    //performs, in sequence, this operation
    //followed by the after operation.
    //This method composes the "before" and
    //"after" operations that we specify in
    //the Consumer<T> object
    //
    //"before" operation is the lambda expression
    //that we first passed to the consumer variable
    //"after" operation is the lambda expression
    //that we passed to after variable
    //
    //once the composition is done, the method
    //will return the Consumer<T> object
    //with the composed operation.
    consumer = consumer.andThen(after);
    
    //Instead of creating "after" variable,
    //we can rely on type inference and directly
    //put the lambda expression in andThen()
    /*
    consumer = consumer.andThen(
    t -> System.out.println(t.toString() + str));
    */
    
    //accept(T t) method accepts "T"
    //type as argument and returns
    //nothing
    consumer.accept("\"");
  }
}
Lambda expression is commonly used in collection. Thus, you will see some prebuilt functional interface being used in Collection API. For example, forEach() method has one parameter which is Consumer<? super T> action. forEach() is a member of Iterable interface, which is implemented by some classes in Collection API like ArrayList, etc.
import java.util.ArrayList;
import java.util.function.Consumer;
public class SampleClass{
  static String str = "";
  
  public static void main(String[]args){
  
    ArrayList<String> arrList = 
    new ArrayList<>();
    arrList.add("A");
    arrList.add("B");
    arrList.add("C");
    
    Consumer<String> consumer = 
    t -> str = str.concat(t.toString());
    
    consumer = consumer.andThen(
    t -> System.out.println(str));
    
    //forEach iterates through 
    //collection where "t" in accept(T t)
    //of Consumer interface is
    //the collection element
    arrList.forEach(consumer);
  }
}
Target Typing in Lambda Expression

Target type is a type in an expression that is expected by the compiler. Target typing is used in typecasting and type inference in generics. Target typing is also implemented in lambda expression.

The Java compiler determines the target type with two other language features: overload resolution and type argument inference. Let's start with type argument inference.
public class SampleClass{
  
  static void concatString(String s,ConcatString cs){
    System.out.println(s.concat(cs.concat("C","D")));
  }
  
  public static void main(String[]args){
    //In this statement the compiler expects that the
    //type in the second argument is ConcatString
    //So, the lambda expression we put in the second
    //argument is going to be ConcatString type
    //
    //The type that is expected by the method is the
    //target type
    SampleClass.concatString("AB",
                            (s1,s2) -> s1.concat(s2));
  }
}

@FunctionalInterface
interface ConcatString{

  String concat(String s1,String s2);
}
For more information about target typing using overload resolution and other information related to target typing, I recommend you to read the The Java™ Tutorials

Lambda Expression as an Alternative to Anonymous Class

One of the usage of lambda expression is to be an alternative to anonymous class. If we want to refer a single abstract method in an interface to a variable then better use lambda.
public class SampleClass{
  
  static void displayString(PrintString ps,
                            String s){
      if(s.startsWith("My"))
         ps.printString(s);
    
  }
  
  public static void main(String[]args){
    
    //anonymous class
    PrintString ps1 = new PrintString(){
    
      @Override
      public void printString(String s){
        System.out.println("Print String: " + s);
      }
    };
    
    //Lambda Expression
    PrintString ps2 = (s) -> System.out.println("String:" + s);
    
    SampleClass.displayString(ps1,"MyString1");
    SampleClass.displayString(ps2,"MyString2");
    
  }
}

@FunctionalInterface
interface PrintString{
  void printString(String s);
}
As we can see, lambda expression is more clear, concise and compact than anonymous class. However, Lambda expression is not a complete alternative to anonymous class.

If we want to refer multiple methods to a variable then, we can use anonymous class but not lambda expression. Lambda expression can be only used to define a single abstract method in an interface.
public class SampleClass{

  public static void main(String[]args){
  
   ClassA a1 = new ClassA(){
   
     @Override
     void meth1(){System.out.println("meth1");}
     
     @Override
     void meth2(){System.out.println("meth2");}
   };
   a1.meth1();
   a1.meth2();
   
  }
}

abstract class ClassA{

  abstract void meth1();
  abstract void meth2();
}
We can also replace anonymous class with lambda expression when implementing prebuilt interfaces with only one abstract method like Runnable, EventListener, etc.

Let's try using lambda expression to implement Runnable interface
public class SampleClass{

  public static void main(String[]args){
    
    //Anonymous Class
    /*
    Thread thread = new Thread(
      new Runnable(){
        
        @Override
        public void run(){
          System.out.println("run!");
        }
     });
     thread.start();
    */
    
    //lambda
    Thread thread = () -> System.out.println("run!");
    thread.start();
  }
}